Artificial Intelligence

AI Chatbot Software Buying Checklist: 25 Requirements

WebPro team 9 min read

A vendor conversation reveals how a company answers hard questions. A requirements checklist is different: a specific list of what the software must actually do, checked against a demo and ideally a pilot — the document that keeps a sales process from skipping something that matters.

A checklist, not a conversation

A vendor conversation reveals how a company answers hard questions. A requirements checklist is different: a specific list of what the software must actually do, checked off against a demo and, ideally, a pilot — the document you hand to whoever runs the evaluation so nothing gets missed under the pressure of a sales process.

This is that list: twenty-five requirements across six categories. It deliberately contains no price ranges or implementation timelines, since both vary enormously by vendor, scope and market and any figure here would be invented rather than researched. Use it alongside, not instead of, the vendor evaluation conversation — this tells you what to check for; that tells you how to read the answers.

Knowledge and accuracy (1-5)

The foundation everything else depends on.

  • 1. Can the platform ingest your actual documentation — existing help articles, PDFs, structured data — without requiring a full manual rewrite?
  • 2. Does every answer show which source document produced it, visible to whoever reviews conversations?
  • 3. Is there a configurable relevance threshold, with a defined behaviour when nothing relevant is found — refusal rather than a forced answer?
  • 4. Can content owners update answers themselves, without a vendor ticket or an engineer?
  • 5. Is there a genuine test environment where you can run your own hardest questions against your own content before committing?

Channels and conversation control (6-10)

Where the bot lives and how much control you have over what it says.

  • 6. Which channels are genuinely supported natively — web, WhatsApp, Instagram, others relevant to your market — versus which require a separate integration effort?
  • 7. Is knowledge, routing and reporting shared across channels, or duplicated and maintained separately per channel?
  • 8. Can you implement a specific business rule — such as escalating after two failed attempts — through configuration, without a development request?
  • 9. Does the platform support the languages your customers actually use, with separate content per language rather than machine translation applied at answer time?
  • 10. Can conversation flows be branched by customer segment, intent or channel, without hard-coding logic outside the platform?

Item 7 is worth probing specifically during a demo — ask to see the same knowledge update reflected across two different channels, live.

Integration and data (11-15)

Where feasibility is actually decided.

  • 11. Can the bot read live data from your systems during a conversation — order status, account details, availability — not only write a record at the end?
  • 12. What integration mechanism is used — native connector, REST API, webhook — and what are the documented rate limits and timeout behaviours?
  • 13. Is every write idempotent, so a network retry cannot create a duplicate record or booking?
  • 14. Where is data processed and stored, in which regions, and is that documented in the contract rather than only in a policy page that can change?
  • 15. Is your conversation data excluded from training any shared or general model, contractually rather than only by assurance?

Items 14 and 15 deserve particular attention — get both in writing in the contract itself, not as a verbal reassurance during the sales process. The integration architecture worth expecting is covered in more depth separately.

Guardrails and permissions (16-19)

What the platform is technically capable of doing, not just what it is instructed not to do.

  • 16. Can specific topics be excluded from the bot's scope at the retrieval layer — so prohibited content is not merely instructed against but structurally unreachable?
  • 17. Are write and action permissions granular and configurable per integration, with the ability to cap amounts or restrict fields?
  • 18. Is there a complete audit log of every action the bot took, with parameters and outcome, retained for a defined period?
  • 19. Can you define a list of actions that always require human confirmation before executing, enforced at the system level?

These map directly onto the permissions model worth expecting from any serious platform — ask specifically whether restrictions are enforced technically or only through prompt instruction, since the difference matters considerably.

Analytics and quality assurance (20-22)

What you can actually see once the bot is live.

  • 20. Does the analytics show resolution rate, escalation rate with reasons, and unanswered questions — not only conversation volume?
  • 21. Can you export raw conversation data for your own analysis, not only vendor-generated summary dashboards?
  • 22. Is there a way to sample and review real conversations against source content, supporting an ongoing quality assurance process rather than a one-time launch check?

Ask to see the actual reporting screen during evaluation, not a screenshot in a sales deck — the gap between the two is common and revealing.

Governance and operations (23-25)

What running this looks like after the project team has moved on.

  • 23. Is there a defined process for content ownership, update approval and change history, or does anyone with access simply edit content directly with no record?
  • 24. What does the exit process look like — can you export your knowledge base, conversation history and configuration in a usable format if you leave?
  • 25. What is required from your team to implement and maintain this — in roles and ongoing effort, not just at launch — stated specifically rather than left implicit?

Item 24 is worth asking in the first meeting, not the last. A platform you cannot leave cleanly is a platform with increasing leverage over you every year you stay.

Scoring the checklist

A structured way to compare vendors against this list, rather than an impression formed across separate conversations.

  1. 1

    Score each item as met, partially met, or not met

    Based on a demonstration, not a claim. 'Yes, we support that' is not a score — seeing it happen is.

  2. 2

    Weight the six categories by what matters most to your use case

    A support-focused deployment weights knowledge and analytics more heavily; a transaction-heavy one weights integration and guardrails more heavily.

  3. 3

    Treat items 5, 14, 15 and 24 as near-disqualifying if unmet

    Testing on your own content, data location and training exclusion in writing, and a clean exit path are foundational enough that a weak answer on any of them should weigh heavily regardless of strength elsewhere.

  4. 4

    Score based on a pilot wherever possible, not a demo alone

    A pilot against real content and a real integration reveals gaps a controlled demonstration does not.

  5. 5

    Involve the people who will operate it daily in the scoring

    Not only the buying committee — support leads and content owners will spot maintenance problems that a procurement-focused review misses.

Score every vendor against the identical twenty-five items, in the identical order, so the comparison is genuinely like for like rather than shaped by which vendor happened to demonstrate their strongest features first.

What the checklist cannot decide for you

This document structures the evaluation. It does not make the decision.

  • Whether your organisation is ready to maintain a chatbot at all — content ownership, ongoing review capacity — regardless of which vendor is chosen.
  • Cultural and tone fit with your brand, which needs a person reading actual sample conversations, not a checklist item.
  • Total cost at your actual volume, which needs modelling with your own numbers rather than a published starting price.
  • Vendor stability and long-term viability, which needs its own separate commercial due diligence.
  • Whether your specific regulatory environment permits what you intend to automate, which needs input from whoever owns that risk in your organisation.

Use the checklist to narrow the field to vendors who pass the technical and operational bar. The final decision among finalists usually comes down to these items the checklist cannot score.

Decision framework and next step

Four questions before starting formal evaluation.

  1. 1

    Have you written your must-have outcomes before speaking to any vendor?

    Score every answer against these, not against how impressive each vendor sounds relative to the others.

  2. 2

    Can you get a pilot on your own content and at least one real integration?

    A demo on sample data tests the vendor's data, not your problem.

  3. 3

    Who from your operations team is involved in scoring, not just buying?

    They catch maintenance problems a purely commercial evaluation misses.

  4. 4

    Are items 5, 14, 15 and 24 non-negotiable for you?

    Decide this before evaluation starts, so a weak answer on a foundational item cannot be argued away by strength elsewhere.

Run this checklist alongside the vendor evaluation conversation — one structures what to check, the other how to read the answers. Our AI solutions overview covers how implementation is typically staged once a platform is chosen.

Frequently asked questions

  1. 1

    What should an AI chatbot buying checklist cover?

    Six categories: knowledge and accuracy, channels and conversation control, integration and data, guardrails and permissions, analytics and quality assurance, and governance and operations — twenty-five specific, demonstrable requirements rather than general impressions.

  2. 2

    How is this different from a vendor evaluation conversation?

    This is a checklist of what the software must demonstrably do, scored against a demo or pilot. A vendor conversation is about how a company answers hard questions about its own limitations — the two are complementary, not interchangeable.

  3. 3

    Which items should be treated as near-disqualifying?

    Testing on your own content before committing, documented data location, contractual exclusion from model training, and a clean, usable exit path. Weakness on any of these is foundational enough to outweigh strength elsewhere.

  4. 4

    Should items be scored from a demo or a pilot?

    A pilot wherever possible. A demo on the vendor's own clean sample data reveals far less than a pilot against your real content and a real integration.

  5. 5

    What can the checklist not tell you?

    Whether your organisation is ready to maintain a chatbot at all, cultural and tone fit, true cost at your actual volume, vendor stability, and whether your regulatory environment permits what you intend to automate.

A checklist structures the evaluation so nothing gets missed under sales pressure. It does not replace judgement about fit, cost and readiness — those still need a person deciding.

Let's talk about your project

Tell us what you want to build and we will work out the scope, timeline and approach together.