
An AI vendor review should test the complete operating dependency, not just model quality. Ask what the product may do, which data it receives, whether that data trains models, how access is controlled, how performance is evidenced, what happens during incidents and model changes, and how the business exits. Score evidence rather than promises, and make high-risk answers contractual before connecting production systems.
Playbook
An AI vendor review should test the complete operating dependency, not just model quality. Ask what the product may do, which data it receives, whether that data trains models, how access is controlled, how performance is evidenced, what happens during incidents and model changes, and how the business exits. Score evidence rather than promises, and make high-risk answers contractual before connecting production systems.
Scope and Authority
Reviewed on 30 July 2026, this questionnaire is educational information, not legal advice. It is a recommended procurement template, not a legal safe harbour, regulator-approved checklist, or substitute for security, privacy, employment, sectoral, and contract review.
The voluntary NIST AI RMF Core calls for policies addressing third-party AI risks and contingency processes for high-risk third-party failures. The NIST Generative AI Profile discusses acquisition due diligence, testing, transparency, and supply-chain considerations.
The NCSC secure-development guidance recommends assessing AI supply chains across the lifecycle and obtaining secure, documented components from verified suppliers. The FTC’s small-business vendor guidance recommends specific security terms, verification rather than trust alone, and rules for vendor data use, sharing, retention, and deletion.
Where a vendor processes personal data, binding data-protection requirements may apply. The ICO processor-contract and due-diligence guidance is UK guidance and should not be treated as universal law, but it illustrates evidence and contractual controls to review with counsel.
How to Score Answers
Score each question:
• 0: Unknown or refused.
• 1: Verbal claim only.
• 2: Documented process or contract statement.
• 3: Tested evidence, independent assurance, or customer-verifiable control.
Mark a separate blocker flag. A high total score should not override a single unacceptable condition such as training on confidential data without authorisation or inability to revoke production access.
45 Due-Diligence Questions
A. Intended Use and Accountability
1. What exact business outcomes and use cases is the product designed to support?
2. Which uses, decisions, industries, or affected groups does the vendor prohibit or caution against?
3. Which party is responsible for configuration, monitoring, human review, and final actions?
4. Can the vendor provide current system, model, data, and limitation documentation?
5. Who is the named escalation owner for product risk, security, privacy, and incidents?
B. Data Collection and Use
6. Which prompts, files, records, metadata, outputs, feedback, and logs does the service collect?
7. Does the vendor or any subprocessor use customer data or outputs to train, fine-tune, evaluate, or improve models?
8. Can all such secondary use be disabled contractually and technically?
9. Where is each data category stored and processed, and which international transfers occur?
10. What are the retention, backup, deletion, export, and verified-deletion procedures for every data category?
The FTC’s AI privacy guidance warns model-service companies to honour privacy and confidentiality promises, including statements about using customer data for model training.
C. Privacy and Affected-Person Rights
11. Does the vendor act as processor, controller, joint controller, or another role for each processing purpose?
12. Which subprocessors receive customer or personal data, and how are customers notified before changes?
13. How does the service support access, deletion, correction, restriction, objection, and portability requests where applicable?
14. What privacy-impact, legitimate-interest, automated-decision, or equivalent assessments has the vendor completed?
15. Can the customer configure data minimisation, redaction, regional processing, and sensitive-data restrictions?
Do not accept “GDPR compliant” as a complete answer. Ask for role, purpose, data, location, contract, and control evidence.
D. Security and Identity
16. Does the service support SSO, MFA, role-based access, and timely user deprovisioning?
17. Do agents and service processes use distinct identities and task-scoped credentials?
18. How are secrets stored, rotated, revoked, and prevented from entering prompts or logs?
19. Which encryption, tenant-isolation, vulnerability-management, penetration-testing, and secure-development controls apply?
20. Which security logs can the customer access, export, retain, and integrate with incident monitoring?
Independent assurance reports help, but scope and exceptions matter. Confirm that the reviewed service, region, controls, and period match the proposed use.
E. Models, Prompts, and Knowledge
21. Which model providers and model versions can process customer data?
22. Can the customer pin, test, approve, or reject a model change before production use?
23. How are system prompts, policies, retrieval sources, and configuration versioned and protected?
24. How does the product detect stale, conflicting, poisoned, or unauthorised knowledge sources?
25. What known limitations exist for factual grounding, supported languages, data formats, and the proposed domain?
F. Agent Actions and Tool Security
26. Which tools, APIs, data stores, networks, and external destinations can an agent access?
27. Can access be limited by operation, object, field, user, tenant, amount, recipient, and time?
28. Which actions require human approval, and is approval technically enforced before execution?
29. How does the service prevent prompt injection, tool poisoning, memory poisoning, loops, duplicate actions, and unauthorised delegation?
30. Are tool calls idempotent, logged with before-and-after state, and reversible or reconcilable?
Map these answers to the official OWASP Top 10 for Agentic Applications when the service can act autonomously.
G. Evaluation, Monitoring, and Change
31. What task-specific evaluation results support performance claims for the proposed use?
32. Can the customer inspect test methods, sample sizes, failure definitions, and known limitations?
33. Which production metrics cover quality, grounding, tool use, policy, cost, human intervention, and affected-user impact?
34. How are regressions detected after model, prompt, data, tool, or policy changes?
35. What release notes, notice periods, rollback options, and customer testing environments are provided?
A generic benchmark does not prove performance on the customer’s documents and workflow.
H. Incidents, Resilience, and Support
36. What constitutes a security, privacy, safety, quality, or agent-action incident for the vendor?
37. Within what contractual period will the vendor notify the customer, and what initial facts will it provide?
38. Can the customer immediately pause workflows, revoke agent access, stop queued actions, and export evidence?
39. What availability, backup, disaster-recovery, data-restoration, and dependency-failure tests have been performed?
40. What support and engineering response is available during a high-impact incident?
Require incident support that is useful before the customer’s own legal or contractual clocks expire.
I. Commercial Terms and Exit
41. What is the full lifecycle cost, including models, tokens, tools, storage, logs, support, implementation, review labour, and data export?
42. Which usage, model, feature, and price changes can occur without customer approval?
43. Who owns customer inputs, outputs, configurations, evaluations, feedback, derived data, and custom workflow assets?
44. Can the customer export data, prompts, logs, evaluations, configuration, and workflow state in usable formats?
45. What happens at termination, including access removal, subprocessor deletion, backup expiry, transition support, and written deletion confirmation?
Compare buying a platform, using a partner, and building internally with the AI consultant vs agency vs in-house guide.

Procurement Decision Framework
Gate 1: Use-case fit
Reject the product if the proposed use conflicts with vendor limitations or if the business cannot define an accountable owner.
Gate 2: Data fit
Do not provide production data until data use, roles, locations, subprocessors, retention, deletion, and training terms are acceptable.
Gate 3: Control fit
Test access limitation, human approval, logs, kill switch, and export. Screenshots and sales claims are weaker than a controlled proof.
Gate 4: Evidence fit
Run a representative pilot against the business’s evaluation cases. Record failures and review load.
Gate 5: Exit fit
Confirm that changing vendors will not strand data, workflow state, evidence, or customers.
Evidence and Artifacts
• Completed questionnaire with owner and date.
• Product and use-case scope.
• Data-flow and subprocessor list.
• Security and assurance documents.
• Privacy and impact assessments.
• Model and change documentation.
• Evaluation report and pilot results.
• Contract, DPA, service levels, and incident terms.
• Pricing model and total-cost estimate.
• Exit test and sample export.
• Exceptions, compensating controls, and approval.
• Annual or change-triggered reassessment.
Common Failure Modes
The review starts after integration. The vendor already has data and access.
A certification replaces scope review. The relevant product or control may be excluded.
The buyer accepts “we do not train on your data” without defining outputs, feedback, logs, and subprocessors.
A demo replaces an evaluation. Curated examples hide failure and review cost.
Incident notice says “without undue delay” but has no operational target. The customer cannot plan its response.
The contract covers data return but not workflow configuration or evidence. Exit still requires rebuilding from zero.
SMB Vendor Review Checklist
• [ ] Define one intended use before sending the questionnaire.
• [ ] Name business, technical, security, privacy, and commercial owners.
• [ ] Mark blockers before calculating the score.
• [ ] Require evidence for material claims.
• [ ] Test with representative, minimised data.
• [ ] Verify human approval and access revocation.
• [ ] Review model and subprocessor change terms.
• [ ] Contract incident notice and cooperation.
• [ ] Calculate full cost, including human review.
• [ ] Run a sample export and deletion request.
• [ ] Record exceptions and compensating controls.
• [ ] Reassess annually and after material change.
FAQs
Should every AI vendor answer all 45 questions?
Use all questions as a screening set, then scale evidence depth to risk. A low-risk writing assistant and an agent with CRM write access should not receive identical review effort.
Is SOC 2 enough for AI vendor due diligence?
No. An assurance report can support security review, but it does not by itself answer model changes, data training, agent authority, evaluation quality, human oversight, cost, or exit questions.
What is the most important AI contract term?
There is no single term. Data use, security, incident notice, change control, ownership, service levels, liability, audit rights, and exit controls work together. Obtain legal advice.
How often should vendor due diligence be repeated?
Review at renewal and after material changes to models, subprocessors, data use, permissions, features, incidents, ownership, or service scope.
What should block a purchase?
Examples include unacceptable data reuse, inability to revoke access, no evidence for consequential claims, missing incident cooperation, prohibited use-case mismatch, or no workable exit. Set blockers before vendor scoring.
Get a 20-Minute AI Workflow Audit
AI Operator can turn your use case into a scored vendor review, pilot evidence plan, control gap list, and implementation decision.