TLDR
- Certifications and logo slides do not predict delivery quality. Method does.
- Fifteen questions, scored 0 to 2, across data readiness, scope, cost forecasting, testing, security, handover, and support.
- Score every shortlisted vendor. Below 20 out of 30 is a delivery risk you should price into the decision.
- The single highest-signal question is number 4: how they decide whether Data Cloud is genuinely required.
- Disclosure up front: Clientell is a Salesforce services partner. We would be a vendor in this process, so score us with the same sheet.
Why the Usual Criteria Fail
The standard evaluation runs on Salesforce partner tier, certification counts, and case studies. All three are weak signals for Agentforce work specifically.
Partner tier measures revenue and historical delivery volume across the whole Salesforce portfolio. It says nothing about agent delivery method.
Certification counts measure exam passes. Agentforce is new enough that certification volume mostly measures how quickly a firm pushed its bench through training, not how many production agents they have operated.
Case studies are selection-biased by construction. Nobody publishes the pilot that stalled. Ask instead about a project that went badly and what changed in their method as a result. The quality of that answer tells you more than five success stories.
What actually separates partners is method: how they assess before they scope, how they justify architecture decisions, how they test things that write to your production org, and what they leave behind. Those are askable questions with checkable answers.
Send the fifteen questions to each shortlisted partner in writing before any demo. Score their written answers, then score again after a live session. The gap between written and live answers is itself informative. Two points for a specific, method-level answer. One point for a plausible but generic answer. Zero for deflection, jargon, or "we'll figure that out in discovery".
Section A: Data Readiness Assessment (Questions 1–3)
1. How will you assess our data quality before scoping, and what will you measure?
2 points. Names specific measurements: duplicate rate per object, field population rates on grounding fields, ownership accuracy against the current territory model, picklist consolidation state, automation inventory on in-scope objects. Describes an assessment deliverable that exists before the build is priced.
1 point. Says they will "review your data" as part of discovery without naming metrics.
0 points. Treats data quality as a client responsibility, or assumes it is fine.
Why it matters. Data remediation is the largest and most variable line in an Agentforce project. A partner who has not measured it has not scoped it, and the number they gave you is a placeholder that will be revised upward.
2. Can we buy the assessment separately, without committing to the build?
2 points. Yes, with a fixed price and a written deliverable you own and can share with other vendors.
1 point. Yes, but credited against the build in a way that creates lock-in, or the deliverable is a presentation rather than a document.
0 points. No, assessment is bundled and only available if you commit to delivery.
Why it matters. A partner who will sell assessment standalone is confident their scope will survive comparison. A partner who bundles it has an incentive to reach the build regardless of what the assessment finds. This is the cheapest leverage in the entire process.
3. What have you found in assessments that caused you to recommend a smaller project?
2 points. Gives a concrete example: a use case they talked a client out of, a Data Cloud build they ruled out, a remediation-first sequence they insisted on.
1 point. Says they would do this in principle but has no example ready.
0 points. Cannot recall recommending less work than the client asked for.
Why it matters. Every partner claims to be consultative. Only some have ever reduced their own revenue on the strength of a finding. This question separates them.
Section B: Implementation Scope Definition (Questions 4–6)
4. How will you determine whether we actually need Data Cloud or Data 360?
2 points. Describes evaluating grounding paths against the use case: record context and prompt template merge fields, Flow actions, Apex actions, then retrieval only if the agent must answer from unstructured content or resolve identity across systems with no shared key. Commits to justifying the recommendation in writing.
1 point. Says it depends on the use case without describing the decision criteria.
0 points. Treats Data Cloud as a mandatory prerequisite for Agentforce.
Why it matters. This is the highest-signal question on the sheet. A default "yes" can add six figures to a project that did not need it. The reasoning behind the correct answer is laid out in does Agentforce require Data Cloud. Ask this question, then check their answer against it.
5. How many agents will you build in year one, and in what order?
2 points. Recommends one agent, shipped properly, with a sequenced roadmap afterwards. Explains that agents two and three cost substantially less once patterns are established.
1 point. Proposes two or three in parallel with a reasonable rationale.
0 points. Proposes a broad multi-agent rollout in year one, or matches whatever number you asked for without pushback.
Why it matters. Parallel first-year breadth is the most common way Agentforce budgets overrun. A partner who agrees to five use cases without argument is optimising for contract value.
6. What is explicitly out of scope, and how are change requests priced?
2 points. Provides a written out-of-scope list and a change-control process with named rates and an approval path.
1 point. Has a change process but no out-of-scope list.
0 points. Vague on both, or claims everything is in scope.
Why it matters. "Everything is included" means the exclusions surface as invoices later. A precise out-of-scope list is a sign of a partner who has delivered enough of these to know where the boundary sits.
Section C: Cost and Consumption Forecasting (Questions 7–8)
7. How will you forecast our Agentforce consumption, and what is your margin of error?
2 points. Describes a method: expected conversation volume, actions per conversation, retrieval calls, retry behaviour, and a stated confidence range. Commits to instrumenting monitoring and alerting before launch.
1 point. Offers a number without showing the derivation.
0 points. Says consumption is a Salesforce matter, or has not considered it.
Why it matters. Ungoverned consumption cancels otherwise successful pilots. A partner who does not forecast it is leaving you exposed to a surprise on a line they influenced through their design choices. Rate structure is in the Agentforce pricing guide.
8. How do you price remediation versus build, and why?
2 points. Separates them. Fixed price for build where scope is known, time and materials with a cap and weekly reporting for remediation where it is genuinely not.
1 point. Fixed price for everything, with a risk premium they can explain.
0 points. Fixed price for everything with no explanation, or time and materials for everything with no cap.
Why it matters. Remediation scope is unknowable before assessment. Fixed-pricing it means you pay a risk premium; uncapped time and materials means you carry all the risk. The honest structure splits it. Cost benchmarks are in Agentforce implementation cost in 2026.
Section D: Testing and QA for Agent Actions (Questions 9–11)
9. How do you test agent actions that write to production data?
2 points. Describes a scenario library with known-correct outcomes, covering ordinary, ambiguous, and adversarial inputs. Tests every action for missing and malformed input. Tests the downstream automation cascade, not just the immediate write.
1 point. Tests conversation quality and does some action testing without a structured library.
0 points. Relies on user acceptance testing or on Salesforce Testing Center alone.
Why it matters. Answer quality and action correctness are different problems. An agent that answers well and acts wrongly 5% of the time is quietly corrupting your data. Testing Center is useful but not sufficient, as covered in Agentforce Testing Center vs real testing.
10. What is your rollback design for agent-initiated changes?
2 points. Snapshot before bulk writes, set-level reversal, a kill switch triggerable by a named client-side person without a deployment, and a rollback procedure executed in a sandbox before go-live.
1 point. Has a documented rollback plan that has not been tested.
0 points. Relies on field history tracking, or has not considered it.
Why it matters. A rollback plan nobody has executed is a hypothesis. This is also the question your risk committee will ask, so it is better to have the answer before the review than during it.
11. How do you handle regression testing when prompts or topics change?
2 points. Treats prompt and topic changes as changes requiring regression against the scenario library, with the library handed to you so you can run it after they leave.
1 point. Regression tests during their engagement only.
0 points. No regression process; prompt changes are treated as configuration.
Why it matters. A prompt edit is a behaviour change with no compiler to catch the consequences. Orgs without regression testing discover breakage through users.
Section E: Security and Permission Review (Questions 12–13)
12. How will you review our sharing and permission model before the agent goes live?
2 points. Maps effective access rather than intended access, reviews field-level security on all in-scope objects, documents the agent's running permission context, and tests with a deliberately low-privilege persona.
1 point. Reviews permissions at object level without persona testing.
0 points. Assumes existing permissions are correct, or treats it as the client's responsibility.
Why it matters. An agent operating in a broader context than the requesting user is a data exposure path with a friendly interface. This is the failure mode that ends programmes rather than delaying them. Background in the Salesforce permissions audit guide.
13. How do you keep sensitive fields out of agent grounding?
2 points. Excludes them by configuration and verifies by testing, rather than instructing the model not to use them.
1 point. Uses a mix of configuration and prompt instruction.
0 points. Relies on prompt instructions alone.
Why it matters. Prompt-level instructions are guidance, not enforcement. Anything that must not be disclosed should be unreachable, not merely discouraged. If you want a fuller security question set, we published one: Salesforce AI agent security vendor questionnaire.
Section F: Handover and Documentation (Question 14)
14. What exactly do we own when you leave, and how will our admins maintain it?
2 points. Names deliverables with acceptance criteria: architecture documentation, action inventory, the scenario test library, runbooks for common failures, rollback procedure, consumption monitoring setup, and structured admin training. Your team demonstrates a change during the engagement, not after it.
1 point. Provides documentation but no training, or training with no documentation.
0 points. Handover is a final-week activity with no named deliverables.
Why it matters. Without proper handover you have bought a dependency. Every prompt tweak becomes a ticket at day rates, and over three years that exceeds what handover would have cost several times over.
Section G: Ongoing Support Model (Question 15)
15. What does support look like after go-live, what does it cost, and what is our internal load?
2 points. Gives a specific model with response times and pricing, and an honest estimate of the internal admin capacity you will need. Distinguishes between what they operate and what you own.
1 point. Offers a support retainer without estimating your internal load.
0 points. Support is undefined, or they claim the agent needs no ongoing management.
Why it matters. Agents are operated systems, not delivered artefacts. A partner claiming otherwise is either inexperienced or selling. Sizing guidance is in the implementation cost guide, and managed services is the alternative to hiring the capacity.
The Scorecard
| # | Question | Area | Score (0–2) |
|---|---|---|---|
| 1 | Data quality assessment method and metrics | Data readiness | |
| 2 | Assessment available standalone | Data readiness | |
| 3 | Example of recommending a smaller project | Data readiness | |
| 4 | How Data Cloud necessity is determined | Scope | |
| 5 | Agent count and sequencing in year one | Scope | |
| 6 | Out-of-scope list and change control | Scope | |
| 7 | Consumption forecasting method | Cost | |
| 8 | Remediation vs build pricing structure | Cost | |
| 9 | Testing method for agent actions | Testing | |
| 10 | Rollback design, tested | Testing | |
| 11 | Regression testing on prompt changes | Testing | |
| 12 | Sharing and permission model review | Security | |
| 13 | Sensitive field exclusion method | Security | |
| 14 | Handover deliverables and training | Handover | |
| 15 | Support model and internal load estimate | Support | |
| Total | /30 |
Interpreting the total
- 26–30. Strong. Method is mature and they have operated agents in production. Compare on price and cultural fit.
- 20–25. Workable. Identify the low-scoring areas and either negotiate them into scope explicitly or plan to cover them yourself.
- 14–19. Material delivery risk. Likely to produce the failure patterns in why Agentforce pilots fail. Proceed only with strong internal capability.
- Below 14. Do not proceed. The gaps are structural, and no amount of contract language compensates for absent method.
If you only weight three questions, make them 4 (Data Cloud justification), 10 (tested rollback), and 12 (permission model review). Question 4 protects the budget, and questions 10 and 12 protect against the two failure modes that end programmes rather than merely delaying them.
Red Flags Outside the Scorecard
A pitch built on shrinking your admin headcount. Beyond being a poor way to treat the people who know your org, it signals a partner who has not thought about who approves agent actions, who maintains grounding content, and who handles escalations. Every production agent has humans operating it. A vendor who says otherwise is describing a governance gap.
No pushback on your scope. If you asked for five agents and they proposed five agents, they are optimising for contract value.
Data Cloud in the quote before the assessment. Architecture decided before evidence is a template, not a scope.
Certification counts as the primary credential. Ask how many agents they have in production and who operates them now.
Reluctance to name a reference doing operations. Implementation references are easy. A reference who has been operating the agent for six months is the one that tells you whether the handover was real.
Disclosure and How We Answer This
Clientell is an AI-led Salesforce services team, so we would be one of the vendors in this process. Score us on the same sheet as everyone else.
For transparency, our answers to the three heaviest questions:
Question 4. We evaluate grounding paths explicitly and put the recommendation in writing, including the reasoning for ruling options out. In practice most first internal agents do not need a Data 360 build, and we say so even though it reduces the project size. Our full reasoning is public in does Agentforce require Data Cloud, so you can check whether our recommendation matches our stated method.
Question 10. Snapshot before bulk writes, set-level reversal, and a client-triggerable kill switch are standard scope rather than options, and the rollback procedure is executed in a sandbox before go-live.
Question 12. We map effective access rather than intended access, and test against a deliberately low-privilege persona before launch. Where an org is still on profiles rather than permission sets, we usually recommend addressing that first.
On question 2, our assessment is available standalone at a fixed price, and the findings document is yours to take to any vendor. That is deliberate. If our scope does not survive comparison, you should know before you sign rather than after.
The one place our model genuinely differs: our agent capability is pointed at the remediation backlog, the high-volume duplicate, ownership, and field-population work that normally consumes the largest share of the budget. Your admin approves batches before anything is written. That compresses the most expensive line in the project. It does not reduce the number of admins you need, and we do not price it as though it does.
Start with the Agentforce readiness audit if you want the assessment first, or book a demo to scope a specific project. If you are still comparing the market broadly, our roundup of top Salesforce consulting companies covers the wider field, including firms we compete with.
Frequently Asked Questions
What should I look for in an Agentforce implementation partner? Method over credentials. Specifically: how they assess data quality before scoping, how they justify whether Data Cloud is needed, how they test actions that write to production, how rollback works, and what you own at handover.
How many partners should I shortlist? Three is usually right. Send the fifteen questions in writing to all three before any demo, and score the written answers before you are influenced by a presentation.
Should I use a large systems integrator or a boutique? Large integrators suit multi-region programmes with formal governance requirements. Boutiques usually give you more senior attention per pound on a single-org project. Score both on the same sheet; the score matters more than the size.
Is Salesforce partner tier a good signal for Agentforce work? Weak. Tier reflects overall revenue and historical delivery volume, not agent delivery method. Ask how many agents they have in production and who operates them today.
Should I pay for a separate assessment? Usually yes. A standalone assessment gives you an unbiased scope and a document you can competitively tender, and it is the cheapest leverage available in the process.
What is a reasonable score to proceed on? Twenty or above out of thirty, with no zeros in the security or testing sections. A zero in those areas is disqualifying regardless of the total.
