What Is a RAG Chatbot? How Retrieval Keeps AI Answers Grounded

What a RAG chatbot is, how one works step by step from crawl to answer, where they fail, how RAG compares with fine-tuning, and when to build one yourself.

Helpo Team · 8 min read
A magnifying glass picking a highlighted passage out of a document and passing it to a chat bubble

A RAG chatbot is a chatbot that looks up relevant passages in your own content, such as docs, a help centre or uploaded files, and passes them to a language model, which writes the answer from those passages. RAG stands for retrieval-augmented generation: the model's generation is augmented with what a search retrieved.

That one change is the difference between a bot that guesses and a bot that answers from what your business has actually written. This post explains how a RAG chatbot works, step by step, using the pipeline we run in Helpo AI as the worked example. It also covers where these bots fail, how RAG compares with fine-tuning, and when to build one yourself.

What is a RAG chatbot?

The easiest way to picture it is an exam.

  • A plain language model sits a closed-book exam. It answers from whatever it memorised during training. It's fluent and fast, but it can't know anything that wasn't in its training data, and when it's unsure it tends to make something up.
  • A RAG chatbot sits an open-book exam. Before answering, it finds the pages in your documentation that match the question, and it's told to answer from those pages only.

The model is the same in both cases. What changes is what it's allowed to look at, and what it's instructed to do when the answer isn't there.

Why not just use the model on its own?

Three reasons, all of which matter for customer support.

  1. It doesn't know your business. No public model has read your returns policy, your pricing or your setup guide. It can only answer in general terms, or guess.
  2. Its knowledge is frozen. A model learns up to a cut-off date. Your product changes every week, and retraining a model each time isn't practical.
  3. You can't paste everything into the prompt. Even with long context windows, sending your whole help centre with every message is slow and expensive, and models get worse at finding the relevant line the more text you give them. Retrieval sends only the few passages that matter.

How a RAG chatbot works, step by step

A RAG chatbot has two halves. Indexing happens ahead of time, whenever your content changes. Answering happens every time someone asks a question.

A pipeline of documents being cut into chunks, turned into vectors, searched and reranked before an answer is written

Indexing: preparing your content

1. Collect the sources. Anything that holds answers: your website or help centre (crawled page by page), PDFs and text files, published help articles, question-and-answer pairs typed by hand, and for a store, the product catalogue.

2. Turn everything into clean text. Web pages are converted to clean markdown, with most of the navigation and page clutter removed. PDFs have their text extracted. Anything that isn't readable text, such as a scanned image with no text layer, can't be retrieved later.

3. Split it into chunks. A model can't search a whole page at once, so each document is cut into passages. Helpo splits on paragraph boundaries and packs them into windows of about 2,000 characters (roughly 500 tokens), with about 200 characters of overlap so a sentence that falls across a boundary isn't lost. Chunk size is a real trade-off. Small chunks match precisely but lose their surrounding context, and large chunks keep the context but dilute the match.

4. Turn each chunk into a vector. An embedding model converts each passage into a long list of numbers that captures its meaning. Passages about the same thing end up close together, even when they use different words. "Can I send it back?" lands near "Our returns policy".

5. Store the vectors with labels. Each vector goes into a vector index, along with the passage text and labels saying which business, which project and which source it came from. Those labels matter more than they look, as the next step shows.

6. Keep it fresh. Pages change, so sources are re-crawled on a schedule. Helpo fingerprints each page's text and skips the re-processing when nothing changed. When something did change, the new passages go live before the old ones are removed, so there's never a moment when the bot has nothing to answer from.

Answering: what happens when someone asks

1. Turn the question into a vector. The visitor's question goes through the same embedding model as your content, so the two can be compared.

2. Search, inside a boundary. The index returns the passages closest in meaning to the question. In a product that serves many businesses, this is where multi-tenancy is enforced. In Helpo every search carries a filter for the business and the project, and there's no code path that searches without one. One customer's bot can never retrieve another customer's documents.

3. Rerank. Vector search is fast but approximate. A second model, called a reranker, reads the question and each shortlisted passage together and scores how well they actually match. The top few, about five in Helpo, go forward. This step does a lot of the work of finding the passage that actually answers the question.

4. Build the prompt. The model gets its instructions, the passages, and the conversation so far. The instructions matter as much as the passages: answer only from the context provided, and if the answer isn't there, say so and offer to connect the visitor with the team. Never invent facts.

5. Write and stream the answer. The model writes a reply grounded in the passages, and it's streamed to the chat widget word by word. Our post on conversation Durable Objects covers the real-time side of that.

6. Record how well it went. Each answer stores how strong the best retrieval match was. That number isn't shown to the visitor, but it's how you find out what your content is missing, which comes up again below.

A worked example

Here's the same pipeline on two questions to a store's support bot.

"Can I return something I bought in the sale?""Do you ship to Iceland?"
Search findsThe returns policy section on sale items, plus the general returns pageShipping pages about the US, the UK and the EU. Nothing mentions Iceland
After rerankingThe sale-items paragraph comes out on topWeak matches only
The bot says"Sale items can be returned within 14 days for store credit…" taken from your policy"I couldn't find that in our shipping information. Want me to connect you with the team?"
What you learnNothing to fixA gap in your shipping page, and a question worth answering

The second column is the one people overlook. A good RAG chatbot is valuable for the answers it gives and for being honest about the ones it can't give.

Where RAG chatbots fail, and how to fix them

RAG makes a bot more reliable, but it doesn't make it perfect. These are the failures we see most often, roughly in order of how often they happen.

  • The right passage isn't retrieved. The answer exists, but it's worded so differently from the question that search misses it, or it's buried in a page about something else. The fix is usually the content: one topic per section, headings phrased like questions, and the plain words customers use. We wrote a whole post on preparing your docs for retrieval.
  • Chunks split the meaning apart. A table cut in half, or an FAQ where the question lands in one chunk and the answer in the next. Q&A pairs avoid this because the question and answer are embedded together.
  • The content is out of date. The bot faithfully repeats last year's pricing. Re-crawl on a schedule and retire old pages.
  • Two sources disagree. An old blog post says 30 days and the policy page says 14, and the bot may pick either. Delete or update the stale source.
  • The model drifts from the context. Instructions help, and so does testing with real questions before launch. The model has to be told what to do when the context is empty, or it'll fall back on what it remembers.
  • The question needs live data. "Where's my order?" isn't in any document. RAG answers from content. Live lookups need a separate integration or a human.
A support team filling a missing piece in a knowledge base after spotting an unanswered question

The way to keep on top of all of these is a feedback loop. Helpo's Knowledge gaps report collects the questions the bot handled badly: weak retrieval matches, "I don't know" replies, thumbs-down ratings and conversations a person had to take over. It groups them by topic and suggests the article that would answer each group. Each new article closes a gap for every customer who asks that question afterwards.

RAG vs fine-tuning

Fine-tuning means training a model further on your own examples. People often assume that's how you "teach a bot your business". For support it's usually the wrong tool.

RAGFine-tuning
What it changesWhat the model can look at when answeringHow the model behaves: its style, format and habits
Updating itEdit a doc, re-index it, and the next answer uses the new versionCollect examples and train again
Can you trace an answer to its source?Yes. The system knows which passages it usedNo. The knowledge is mixed into the model's weights
Risk of stale answersLow, if sources are re-syncedHigh. It's frozen at training time
Best forFacts that change: policies, prices, product detailsA consistent tone or output format

The short version: fine-tuning teaches style, RAG supplies facts. Most support bots need RAG first. Many never need fine-tuning at all, because a good system prompt handles tone well enough.

Build one or buy one?

Building a basic RAG chatbot is a well-documented weekend project. Open-source frameworks such as LangChain and LlamaIndex, plus any hosted vector database, will get a demo answering questions over a folder of PDFs in an afternoon.

The demo isn't what takes the time, though. What takes the time is everything a production support bot needs around it:

  • crawling and re-syncing websites without duplicating content
  • strict per-customer isolation in the index
  • reranking and testing answer quality on real questions
  • a chat widget that works on any site, on mobile, in many languages
  • handing a conversation to a human mid-chat, with a shared inbox
  • tickets, lead capture, analytics, and a report of what the bot couldn't answer

Build if the chatbot is your product, if your data can't leave your own systems, or if you need retrieval over unusual data. Buy if what you actually need is customer support, and the bot is a means to that.

Helpo AI is the second kind. It runs the pipeline described above on your website, docs and files, hands off to your team in a shared inbox, and comes with tickets, lead capture and the Knowledge gaps report. If you're weighing ChatGPT for the same job, our guide to ChatGPT for customer service covers where it fits and where it doesn't.

Helpo AI

An AI agent that answers from your own docs

Start free

Frequently asked questions

What does RAG stand for?

Retrieval-augmented generation. The chatbot first retrieves the passages in your content that match the question, then a language model generates the answer from those passages.

Is ChatGPT a RAG chatbot?

Not in a normal conversation, where it answers from what its model learned in training. It does retrieve when it searches the web or reads files you've uploaded, which is the same idea. A RAG support bot does that on every question, over your own content only.

Does RAG stop AI hallucinations?

It cuts them down a lot, but it doesn't eliminate them. The model can still misread a passage or fill a gap. Strict instructions to answer only from the retrieved content, a clear "I don't know" path and regular review of real conversations keep the rest in check.

Should I use RAG or fine-tuning for a support bot?

RAG. Support answers are facts that change, like policies, prices and product details, and RAG picks up an edit as soon as the content is re-indexed. Fine-tuning changes how a model writes, not what it knows, and has to be redone whenever your content changes.

What content works best for a RAG chatbot?

Short, focused help articles with descriptive headings, written in the words your customers use. Q&A pairs work well for one-line answers. Long pages that mix many topics, scanned PDFs and outdated pages work worst.

Share this article
Helpo AI

An AI agent that answers from your own docs

  • Live on your site with one script tag
  • Hands the chat to your team when a human is needed

No card needed · 14 days of Pro

Newsletter

New posts in your inbox

Practical notes on AI support and running a helpdesk. Unsubscribe with one click.