Skip to main content

What is retrieval-augmented generation (RAG)?

Retrieval-augmented generation, or RAG, is a way of answering from a body of documents: the system first finds the passages most relevant to a question, then gives them to the model to answer from.

It is how an AI can answer about your files without reading all of them on every message, and without being retrained on them.

How it works

Documents are split into chunks of a few hundred words. When a question arrives, the system scores the chunks against it, by matching words, by comparing meaning, or both, and takes the best few. Those passages go into the model's context with the question, usually with an instruction to answer from them.

Examples

  • Asking about a contract and getting the clause the answer rests on.
  • Searching a folder of notes by what they say, not their file names.
  • A support assistant that answers from a product manual.

Limitations

  • If retrieval misses the right passage, the model answers without it, or from general knowledge.
  • Chunks lose context: a sentence can depend on a page that was not retrieved.
  • It reflects the documents as they were when they were added.

How Venta uses it

Documents you add are split into chunks, and each message is matched against them. The best passages are sent with your question, and the reply's steps show which documents it read.

If the search of your documents fails, Venta says so near the start of the reply instead of quietly answering from general knowledge.

Questions people ask first

Is RAG the same as training a model on my data?
No. Training changes the model itself. RAG leaves the model as it is and hands it the relevant passages each time.
Does RAG stop AI from making things up?
It reduces it when the right passage is found, but does not prevent it. Check the source for anything that matters.
Which files can Venta read?
Reads PDFs, Word documents, spreadsheets, CSV, HTML and text files you upload.