Designing an AI Concierge That Qualifies Visitors With Facts You Publish

A website concierge answers, asks and recommends from your own catalogue and rules, quotes prices instead of inventing them, and hands the conversation to a person with a summary that can be

Netics editorial card showing a website assistant answering from a catalogue and handing a qualified conversation to a person.
Netics editorial card on a website concierge: answers drawn from published facts, a qualification step, and a hand-off a person can act on.

TL;DR

  • An AI concierge is a customer-facing assistant on your website or a messaging channel that answers questions about your offer from your own content, asks qualifying questions and passes the conversation to a person with a summary.
  • The design question is the boundary: what the assistant may answer from, which questions it must ask, and what it hands over when a visitor is ready to talk.
  • Your knowledge base sets the answer ceiling, because the catalogue, the FAQ, the pricing pages and the rules for what to recommend form the assistant's whole vocabulary.
  • Qualification is the step most contact forms skip. The assistant asks for the constraint that decides the recommendation, such as volume, deadline, site count or team size, and records the answer in the visitor's own words.
  • Pricing stays a citation. The assistant quotes the price on your page, and anything that requires a negotiation leaves the conversation for a person.
  • The hand-off carries four fields: the need, the fit, the boundary of what was promised and the owner who picks it up, which is what turns a chat transcript into a working lead.
  • Cost and abuse limits belong in the first version: message caps per session, rate limits per visitor, bot protection and a spend ceiling tied to the model provider.
  • Built from components you can host, meaning an automation platform, retrieval over your own documents, a model of your choice, your calendar and your email, the assistant keeps visitor conversations inside infrastructure you control.
  • The outcome worth measuring is a faster first response with a summary a human can act on, which is a different metric from chat volume.

What an AI concierge does on a website

Most visitors leave without asking anything. The questions they would have asked arrive later by email, in a form that says "please contact me", or never at all. A concierge exists to move that conversation to the moment the visitor is actually on the page, and to leave behind a record of what they wanted.

The mechanics are unglamorous. A chat surface sits on the site or in a messaging channel. Behind it, an assistant answers from your content, asks the questions your team asks on every sales call, and then writes a short summary that reaches a human through email, a calendar booking or a CRM entry. Netics' own concierge page describes exactly that sequence across four steps, from learning the offer to handing the conversation over.

Netics editorial definition card setting out the answer boundary of a website concierge, covering what it may answer from, what it asks and what it hands over.
Netics editorial card: the boundary is the design. Everything the assistant says comes from published facts, and everything it cannot support becomes a hand-off.

The reason to build it from your own material rather than from a general model is the same reason retrieval exists at all. A language model asked about your business will produce fluent answers that sound right and contain numbers nobody approved. Grounding changes the question from "what does the model know about us" to "what do our documents say", which is a question with a checkable answer.

The knowledge base sets the answer boundary

Everything the assistant can say comes from something you already published or wrote down for it. That is a small idea with large consequences, because it converts a content problem into a maintenance problem.

The sources that matter for a sales conversation are usually already sitting on the site: product or service pages, a pricing page, an FAQ, delivery or lead-time notes, and a short list of rules about what to recommend when a visitor describes a need. Add the internal version of the same knowledge, which is the set of constraints your team applies before quoting: minimum order quantities, supported regions, onboarding effort, what you decline to take on.

Retrieval is what makes that usable at conversation speed, and filtering is what keeps it precise. Qdrant's documentation describes constraining a query by payload filters, which is the mechanism that keeps an answer about a small deployment from quoting terms that apply to an enterprise contract. In practice the filter is built from what the visitor has already said: the product family, the region, the volume band.

Two rules keep the boundary honest as the assistant grows. First, every answer should be traceable to a document, so a wrong answer becomes a content fix rather than an argument about the model. Second, the assistant's vocabulary needs an owner. A pricing page that changes without the knowledge base being re-synced is the most common way these projects go wrong, which is why the sync belongs in the deployment rather than in somebody's calendar reminder.

Capture of n8n's official documentation for the AI Agent node, showing the node description and its requirement for at least one connected tool.
n8n's official documentation for the AI Agent node (docs.n8n.io, retrieved 2026-10-07): the node builds an agent from a chat model plus connected tools, and the documentation states that at least one tool sub-node is required.

Qualification asks the questions your team keeps repeating

Qualification is where a chatbot becomes a sales tool, and it is also where most implementations stay shallow. A form asks for a name, an email and a message. A concierge asks for the one thing that decides the recommendation.

For a service business that is often volume, deadline or scope. For a product business it is usually the deployment shape: how many sites, how many users, which of the two models on the pricing page. The assistant should ask for that constraint before it recommends anything, and it should ask in the visitor's own terms rather than forcing a dropdown.

There are three practical rules worth building in from the start. Ask one question at a time, because a wall of four questions reads as a form. Record the answer verbatim alongside the summary, because paraphrases lose the detail a salesperson would have noticed. And accept "I don't know" as an answer, since a visitor who cannot state a budget is a visitor who needs the explanation more than the quote.

The tooling for this is standard. n8n's documentation describes the AI Agent node as a chat model plus connected tools, where the agent chooses which tool to call, and it requires at least one tool sub-node. The retrieval side has its own node: n8n's Qdrant Vector Store node documentation states that the node can be connected directly to an agent's tool connector, so the lookup against the knowledge base becomes one of those tools. A lookup against the knowledge base, a booking action and a hand-off action are three tools, which is enough structure for a qualifying conversation.

Handing the conversation over with a usable summary

The hand-off is the part that decides whether anyone trusts the assistant. A raw transcript leaves the salesperson to reread the whole conversation; four fields turn it into a lead they can act on.

Netics editorial scorecard listing the four fields a concierge hand-off carries: the need, the fit, the boundary and the owner.
Netics editorial card on the hand-off: four fields turn a transcript into a lead the person who picks it up can act on.

The need comes first, written as the visitor expressed it. The fit follows, naming the offer the assistant recommended and the reason it chose that one. The boundary records what was promised, including the price quoted and the source page it came from, which is what protects both sides in the follow-up call. The owner closes it: who receives it, through which channel, and by when.

Netics' published description of the concierge ends its flow with exactly this step, sending the summary by email, booking the call, or creating the lead in the client's own tools. The detail worth copying is that the assistant stays presented as an assistant. Visitors who know they are talking to an AI ask better questions than visitors who suspect they are talking to a badly supervised human.

Retrieval quality shows up here as well. If the summary says the assistant recommended the wrong service level, the fix is usually a missing document or a filter that was too broad, and both are content edits. Treating those reports as a weekly review item is what keeps a concierge from drifting into confident nonsense after six months of catalogue changes.

Cost and abuse limits belong in the first version

A concierge is a public endpoint attached to a paid model, which makes it a cost centre with an audience. Limits designed after launch tend to arrive as an outage.

Four controls cover most of the risk. A cap on messages per session keeps a looping visitor from running indefinitely. A rate limit per visitor address handles the obvious scripted case. Bot protection sits in front of the chat surface, because a form-protection layer that only guards the contact page leaves the new endpoint open. A spend ceiling with an alert threshold at the model provider turns an unexpected pattern into a notification rather than an invoice.

There is a fifth that is easier to miss: an answer length limit. Long answers cost more, read worse on a phone and give a visitor less chance to correct the assistant's understanding. Short answers with a follow-up question are both cheaper and more useful.

Capture of Qdrant's official filtering documentation, showing payload filter conditions that constrain which retrieved points an answer can use.
Qdrant's official filtering documentation (qdrant.tech/documentation, retrieved 2026-10-07): the payload filter is the mechanism that keeps an answer inside the product family, region or volume band the visitor actually asked about.

Keep the logs too, with a retention window you can defend. Conversation logs are what let you answer the two questions that will come up: which questions do visitors keep asking that our content does not answer, and which answers did the assistant get wrong. For a company with any privacy obligation at all, that log is personal data, so the retention decision belongs with whoever owns the privacy policy.

Building it from components you can host

The stack that makes this practical is deliberately boring, and our own deployment runs on the same pieces: an automation platform for the conversation flow, retrieval over your documents, a model of your choice, and your existing calendar and email.

The choice that matters commercially is where the conversation data lives. A hosted assistant service gets you running in an afternoon and puts your visitor conversations and your pricing rules inside somebody else's account. The hosted-component version takes longer, and it keeps the retrieval index, the logs and the model credentials inside infrastructure you control, with a provider you can change when prices or policies move. Our own write-up of the components behind an AI concierge covers the four steps of the flow, and the automation platform that runs it explains why an agent with tools fits this job better than a fixed command list.

Model choice deserves its own sentence, because it is the part teams over-think. A qualifying conversation with retrieval is a task small and mid-size models handle well, and the model is the easiest component to swap later. Retrieval quality, document freshness and the hand-off format decide whether the assistant is useful, which is why the grounding and refusal behaviour of a private assistant is the part of that article worth reading before choosing a provider.

Netics editorial banner stating that qualification runs on facts the company already publishes, with everything outside those facts becoming a hand-off.
Netics editorial card: the knowledge base is the product surface. Adding a document changes what the assistant can say, and that is the lever a team actually holds.

One operational note closes the loop. Treat the concierge as a product with a release cadence rather than a one-off build. A monthly review of unanswered questions, a quarterly review of the pricing and rules documents, and a check on the cost ceiling covers most of what goes wrong between launches.

If you want a second opinion on whether this fits your sales process, book a free 30-minute audit and bring the questions your team answers most often. That conversation is usually enough to tell whether a concierge would help and where its boundary should sit.

Sources

Source: AI Agent node documentation — docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent/, n8n, retrieved 2026-10-07 (the node builds an agent from a chat model plus one or more connected tools and the agent decides which tools to call, the requirement to connect at least one tool sub-node, the deprecation of the agent type setting in favour of the tools agent pattern). Source: Qdrant Vector Store node documentation — docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.vectorstoreqdrant/, n8n, retrieved 2026-10-08 (the node uses a Qdrant collection as a vector store for inserting, getting and retrieving documents, and it can be connected directly to the tool connector of an AI agent to use the vector store as a resource when answering queries). Source: Filtering — qdrant.tech/documentation/concepts/filtering/, Qdrant, retrieved 2026-10-07 (constraining retrieved points with payload filters built from match conditions, ranges and nested conditions, applied to a query and to scroll or count operations). Source: AI Concierge for Websites — neticslabs.com/en/solutions/concierge, Netics, retrieved 2026-10-07 (the four steps from learning the offer to handing over, the statement that the assistant "never invents prices or promises", the hand-off through email, calendar booking or a CRM entry, the component list of n8n, retrieval over documents in Qdrant, a chosen model, calendar and email, and the inclusion of abuse and cost limits such as bot protection and message caps from day one). Figures: the n8n documentation capture from docs.n8n.io and the Qdrant documentation capture from qdrant.tech, plus three Netics editorial diagrams rendered from the same sources.