Reference architectures
RAG architecture
Retrieval-augmented generation in its basic form: ingest documents, retrieve context, prompt the model.

A RAG system has two paths. The ingestion path loads documents, splits them into chunks, turns each chunk into a vector with an embedding model and stores it in a vector database. The query path takes the user's question, retrieves the most similar chunks and sends them to the LLM together with the question. The diagram shows both paths meeting at the vector database. Use it as the starting point for any assistant that has to answer from your own documents.


