Artificial Intelligence

B2B Chatbot Buying Guide: 12 Questions to Ask Before Choosing a Platform

WebPro team 11 min read

Chatbot demos hide exactly the things that decide whether a deployment works: knowledge maintenance, escalation, integration limits and what happens to your data. This is the evaluation conversation rather than a feature matrix — twelve questions, why each matters, and how to tell a specific answer from a reassuring one.

What goes wrong in chatbot procurement

Chatbot demos are unusually good at hiding the things that decide whether a deployment succeeds. The demo runs on a clean dataset, answers questions the vendor chose, and never has to hand a frustrated customer to a support agent at 11pm. Six months later the same platform is failing on the parts nobody asked about: knowledge maintenance, escalation, integration limits and what happens to your data.

The remedy is not a longer feature comparison. Feature matrices reward vendors who tick boxes and penalise ones who answer honestly. What separates platforms in production is how they behave under conditions the demo avoids — and the only way to find that out is to ask specific questions and listen to how they are answered.

This guide is the evaluation conversation: twelve questions, why each one matters, and what a weak answer sounds like. It is deliberately vendor-neutral and it is not a requirements document — the requirements come after you have worked out which vendors can have this conversation credibly at all.

Questions about knowledge and accuracy

Most chatbot disappointment traces back to knowledge rather than to the model. Start here.

  1. 1

    Where do answers come from, and can you show me the source for a given reply?

    A platform that cannot show which document produced an answer cannot be debugged. When answers go wrong — and they will — the first question is always 'why did it say that?'. A weak answer sounds like 'the AI learns from your content'. A strong one shows you a citation and a retrieval trace.

  2. 2

    How does content get updated, and who can do it?

    If updating an answer requires a vendor ticket or an engineer, your knowledge will go stale. Ask to see the workflow a non-technical subject matter expert would use, end to end.

  3. 3

    What happens when the bot does not know?

    The honest behaviour is to say so and escalate. Ask to see it. Platforms that always produce a confident answer are producing confident wrong answers some of the time, which is worse than silence in a B2B context.

  4. 4

    How do you prevent answers outside approved content?

    Ask specifically about guardrails: topic restrictions, refusal behaviour, and whether the bot can be prevented from discussing pricing, legal terms or competitors. 'The model is very accurate' is not an answer to this question.

Questions about channels and conversation design

The second cluster is about where the bot lives and how much control you have over what it says.

  1. 1

    Which channels are genuinely supported, and is it one bot or several?

    Many platforms support WhatsApp, Instagram, web chat and phone as separate products with separate configuration. Ask whether knowledge, routing and reporting are shared across channels or duplicated per channel — the maintenance cost of the second model is significant.

  2. 2

    How much control do we have over the conversation flow?

    Ask how you would implement a specific rule — for example, always escalate after two failed attempts, or never discuss pricing above a published range. If the answer is 'we can configure that for you', find out what happens when you want to change it on a Friday.

  3. 3

    What does multilingual support actually mean here?

    Ask whether each language has its own knowledge base or whether one is machine-translated at answer time, and how quality is measured per language. This matters more in markets like ours, where a business often serves customers in Azerbaijani, Russian and English at once.

  4. 4

    How does handover to a person work?

    Ask to see what the agent receives: full transcript, resolved customer identity, detected intent, escalation reason. Ask what happens outside business hours. Weak platforms treat escalation as an afterthought, which is exactly where routing rules tend to break.

Platforms that run one configuration across web, messaging and phone — the approach taken by tools like Vexvon — remove a class of maintenance problem, but they should still answer every question above individually rather than pointing at a diagram.

Questions about integration and data

This is where feasibility is decided, and where vague answers cost the most.

  1. 1

    How does the platform read from and write to our systems?

    Ask for the mechanism, not the logo wall. Native connector, REST API, webhook, middleware? What are the rate limits? What happens to a conversation when the CRM is slow or down? A logo on a slide is not an integration.

  2. 2

    Can it call our systems during a conversation, not just after it?

    Checking an order status, a booking or an account balance mid-conversation is a different capability from writing a lead record at the end. Many platforms only do the latter. Be explicit about which you need.

  3. 3

    Where is our data stored and processed, and for how long?

    Ask for the regions, the sub-processors, the retention period and the deletion mechanism, in writing. Whether that arrangement is acceptable is a decision for whoever owns data protection in your organisation — get them the written answer rather than a verbal assurance.

  4. 4

    Is our conversation data used to train shared models?

    Ask directly, and ask for the contractual position rather than the marketing position. The answer may be perfectly acceptable; what matters is that it is documented.

If the answers here are thin and the platform is otherwise strong, the gap is usually bridgeable with integration work on your side. Our custom software services page covers how that layer is normally built when a platform stops short of what a business needs.

Questions about operations and cost

The last cluster is about what running this looks like after the project team has moved on.

  1. 1

    What does the analytics actually show?

    Ask to see the reporting screen, not a screenshot. You need resolution rate, escalation rate with reasons, unanswered questions and per-channel breakdowns. Conversation volume alone tells you nothing about whether the deployment is working.

  2. 2

    How is pricing structured as we grow?

    Per conversation, per resolution, per seat, per channel? Model your expected volume at two and five times current levels and ask for the figure. Pricing that is comfortable at pilot scale sometimes is not at production scale, and that is better discovered now.

  3. 3

    What does implementation require from us?

    Ask for the list of what they need from your team, in hours and in roles. Underestimating this is the single most common reason chatbot projects slip — the work is usually content preparation and integration, both of which land on your side.

  4. 4

    What does exit look like?

    Can you export your knowledge base, conversation history and configuration in a usable format? A platform you cannot leave is a platform with increasing leverage over you every year.

How to run the evaluation itself

The questions matter less than the conditions you ask them under. A structured evaluation surfaces differences that a demo cycle hides.

  • Write down your three or four must-have outcomes before speaking to any vendor, and judge every answer against them rather than against each other.
  • Run the same scripted scenario with every vendor, including one deliberately difficult case — an angry customer, an ambiguous question, an out-of-scope request.
  • Insist on a pilot with your own content and at least one real integration. A pilot on sample data measures the vendor's data, not your problem.
  • Involve the people who will operate it daily — support leads, not only the buying committee. They will spot maintenance problems buyers miss.
  • Score answers on specificity, not enthusiasm. 'Yes, via webhook, with a 30-second timeout' outranks 'absolutely, that is fully supported'.
  • Keep a written record of commitments made verbally and have them reflected in the contract.

Two weeks of structured evaluation is cheap compared with a year on a platform that cannot escalate properly.

Metrics to agree before you sign

Agree how success will be judged before procurement closes, while you still have leverage and while expectations are still adjustable.

  1. 1

    Containment or resolution rate, defined precisely

    'Resolved' has to mean the customer got what they needed, not that the conversation ended. Agree the definition in writing, because vendors and buyers routinely mean different things by it.

  2. 2

    Escalation rate with reasons

    Not just how often the bot hands over, but why. This is the number that tells you what to improve, and it should be available without a support request.

  3. 3

    Answer accuracy on a sampled review

    A weekly human review of a random sample of conversations. There is no automated substitute for this, and platforms that cannot support the sampling workflow make it harder than it needs to be.

  4. 4

    Time to update an answer

    From 'we noticed a wrong answer' to 'it is corrected in production'. If that is measured in days, knowledge will always lag reality.

Review these by channel as well as in aggregate — a platform can perform well on web chat and poorly on messaging while the average looks acceptable.

Where a platform should stop and your team takes over

No platform decision removes the need for human ownership. Be clear during evaluation about which responsibilities stay with you.

  • Deciding what the bot is allowed to say about price, contracts and commitments — that is a business decision, not a configuration default.
  • Owning the knowledge base. Vendors can host it; someone in your organisation has to be accountable for whether it is right.
  • Reviewing conversations regularly. Automated quality scores are useful signals, not a substitute for reading what customers are actually asking.
  • Handling complaints, legal matters and anything involving regulators.
  • Deciding when the bot's scope expands. Scope creep by default is how accuracy quietly degrades.

A vendor who is candid about these boundaries during a sales process is usually a better long-term partner than one who implies the platform handles everything.

Decision framework and next step

Reduce the shortlist with four questions, in this order.

  1. 1

    Can it answer accurately from our content?

    Tested on your material, not theirs. If this fails, nothing else matters.

  2. 2

    Can it reach our systems in the way we need, during a conversation?

    Read and write, not just write. Confirm with a working pilot integration.

  3. 3

    Can our team maintain it without the vendor?

    Content updates, rule changes, escalation configuration. If not, budget for the dependency explicitly.

  4. 4

    Does the commercial model still work at three times our current volume?

    Model it before signing, not after growth makes it urgent.

If two platforms pass all four, choose the one whose escalation and reporting your operations team prefers — that is what they will live with daily. For how the conversation itself should be designed once a platform is chosen, see chatbot lead qualification; our AI solutions overview covers how implementation projects are usually staged.

Frequently asked questions

  1. 1

    What should a B2B chatbot buying guide actually cover?

    The conditions a demo avoids: where answers come from, how knowledge is maintained and by whom, how escalation works, what integration is really possible during a conversation, where data is processed, and what the commercial model does as volume grows.

  2. 2

    What integrations should we insist on before signing?

    At minimum, real-time read access to whichever system holds the answers customers ask about, write access to your CRM with structured fields, and a tested escalation path into the tool your agents already use.

  3. 3

    Which metrics should be agreed with the vendor up front?

    A precise definition of resolution, escalation rate with reasons, a sampled human accuracy review, and the time it takes to correct a wrong answer in production.

  4. 4

    What are the biggest procurement mistakes?

    Judging on a demo run with vendor data, treating a logo wall as proof of integration, leaving the definition of 'resolved' vague, and not modelling cost at realistic future volume.

  5. 5

    What stays the buyer's responsibility regardless of platform?

    Owning the knowledge base, deciding what the bot may say about commercial terms, reviewing real conversations, and handling complaints and legal matters.

If you are still deciding whether to buy a platform at all rather than build on top of existing systems, that is a separate question with a different decision framework.

Let's talk about your project

Tell us what you want to build and we will work out the scope, timeline and approach together.