Understanding RAG from Scratch (II): Running a Local RAG Demo on Mac mini—A Practical Guide to the Minimal Architecture
本文最后更新于 218 天前,其中的信息可能已经有所发展或是发生改变,如有失效可到评论区留言。
Article Abstract
Addressing the need for a fully localized blog chatbot, this article uses a Mac mini as the platform to build a minimal RAG Demo using Python, analyzing the complete workflow from document splitting and vectorization to vector retrieval. It uses Hugging Face's transformers library to load the embedding model and LLM, combined with FAISS to build a vector index, achieving local closed-loop validation of the knowledge base. This practice covers the technical details of the core RAG modules, providing a foundation for subsequent formal deployment based on PVE LXC, while serving as an advanced chapter in the tutorial series to help readers deeply understand the construction logic of local RAG systems through reproducible code examples after mastering the principles.
Qwen3-14B · 2026-06-18

1 Introduction

In previous articles (see article: Home Data Center Series - Understanding RAG from Scratch (1): Principles and Complete Workflow Analysis), I introduced the theoretical five-step workflow of RAG: “Splitting → Vectorization → Vector Storage → Retrieval → Answer Generation”; subsequently, based on the self-built embedding model on Ollama nomic-embed-text, I implemented a knowledge base on Chatbox, which can be considered the simplest RAG practice (see article:Home Data Center Series: Building Your Own Embedding Model with Ollama + Chatbox Knowledge Base Hands-on Practice)。

However, this implementation method of the Chatbox knowledge base actually hands over the three steps of “splitting, vector storage, and retrieval” to Chatbox, while the answer generation part is executed by the large language model specified by the user (I used gpt-5-mini in the article). For daily personal use, this method is sufficient, but if you want to make a completely free blog chatbot, this method won't work.

To truly achieve a fully local, free blog chatbot, you must build a complete local RAG workflow yourself. In theory, this can be done with the help of existing frameworks, such as LangChain 或 LlamaIndex to implement. However, in order to first get familiar with the entire RAG process, I decided to first build a Python-based minimal Demo, using transformers to call Hugging Face's embedding models and LLM models: not aiming for complex features or performance optimization, but simply to run through the complete closed loop of “document → embedding → retrieval → calling LLM to generate answers”. Compared to Ollama or other inference platforms, this approach is more lightweight, controllable, and fully localized, making it perfect for beginners to understand every step of the RAG process.

Through this Demo, I hope to truly understand the internal logic and data flow of each step, laying a solid foundation for building a formal blog chatbot in PVE LXC later. In other words, this article is astep-by-step demonstration of a minimal local RAG processlearning case study, allowing you to get familiar with the RAG process first before considering using frameworks or deploying to a production environment.

Note: The goal of the simplified Demo is “lightweight, controllable, and fully local”, suitable for learning the RAG process; whereas Ollama or other inference platforms are oriented towardsOfficial Deployment, supporting larger models, higher performance, and online services. During the learning phase, you can absolutely use this minimal Demo to understand the process first, and then gradually upgrade to a formal environment.

2 Preparation

1 Preparing the Hardware Environment

To run a minimal local RAG process, the hardware threshold is actually not high. My practical environment this time was completed on an M4 Pro Mac mini (24GB RAM) , which is configured well enough to smoothly run small Hugging Face models (such as sentence-transformers/all-MiniLM-L6-v2 as the embedding model, and Jackrong/llama-3.2-3B-Chinese-Elite-v2 as the LLM).

Of course, you don't necessarily need the same Mac mini as mine; as long as your computer meets the following conditions, you can basically run it:

  • CPU supports AVX2 instruction set(most Intel/AMD processors from recent years are fine; Apple silicon also natively supports it under macOS).
  • Memory at least 16GB(running a 3B scale model is safer; if you want to run an 8B model, 24GB+ is recommended).
  • Disk space over 15GB(mainly used to store Hugging Face model files and vector database data).
  • Internet access(internet is required for the first-time model download; if the download is interrupted, you can also manually download the model files and place them in the local cache directory; subsequent inference runs completely locally).

In other words, as long as you can install Hugging Face's dependency libraries on your device and run a model of appropriate size, you can complete the minimal RAG Demo shown in this article.

2 Preparing the Software Environment

2.1 Python Environment

To run the minimal RAG Demo, you first need to prepare a suitable Python environment.

The Python that comes with macOS is usually 3.9:

image.png

which is too old and incompatible with the dependencies we will use later, so a separate installation of version 3.11 or aboveof Python is required. On macOS, there are two main ways:

Method 1: Install via Homebrew (Recommended)

This is the most common and recommended way for macOS users, as Homebrew isolates well from the system, does not affect the built-in Python, and makes future upgrades or uninstallation convenient.

If you just want stability, it is recommended to directly install Python 3.11:

brew install [email protected]

After successful installation, the displayed result is similar to the following:

image.png

You can also confirm with the following command:

python3.11 --version

image.png

If you want to try the latest, you can also directly install the latest release version provided by Homebrew (currently 3.14):

brew install python
python3 --version

python3 --version

Tip: Although the latest release version of Python has reached 3.14, some dependencies may not have fully caught up yet. To reduce compatibility issues, installing 3.11 is highly recommended as a safe choice.

Method 2: Use the official Python website installer (pkg)https://www.python.org/downloads/The official Python website (

image.png

) also provides a macOS installer (.pkg format). The installation method is very intuitive—just download and double-click to complete:

Therefore,This method installs Python to the system path (/Library/Frameworks), which is less elegant than Homebrew management and may conflict with the system's built-in version.; while macOS users are recommended to use HomebrewNon-macOS users (such as Windows, Linux)

2.2 pip and Dependency Installation

2.2.1 Confirming the pip version bundled with Python 3.11 installation

can download the installer for their corresponding platform from the official Python website.

python3.11 -m pip --version

python3.11 -m pip --version

image.png

2.2.2 Manually installing or upgrading pip as needed (Optional)

If you find that pip is missing, you can install it manually:

curl -sS https://bootstrap.pypa.io/get-pip.py | python3.11

It is recommended to upgrade pip to the latest version to avoid errors in subsequent dependency installations:

python3.11 -m pip install --upgrade pip

If successful, the output is as follows:

image.png

2.2.2.3 Install Dependencies

To run the minimal RAG Demo, we need some basic packages:

python3.11 -m pip install -U faiss-cpu numpy scipy torch torchvision sentence-transformers transformers tqdm safetensors

Description:

  • faiss-cpu: An efficient vector search library (CPU version, Mac mini is sufficient).
  • numpy: A scientific computing library that libraries like faiss depend on.
  • scipy: A common scientific computing library used for similarity calculations (such as cosine similarity).
  • sentence-transformers: A common text vectorization toolkit used to convert split text chunks into vectors.
  • transformers: Hugging Face's model library, which can call Embedding models or simple LLMs.
  • tqdm: A progress bar tool that displays real-time progress when batch processing split text chunks, making it more intuitive.

The result after installation is as follows:

image.png

In this way, the Python environment has the minimal dependencies to run the RAG Demo.

Note: The dependencies installed here are only the most basic packages required to run the minimal RAG Demo. In the subsequent use LangChain 或 LlamaIndex to build a formal local RAG system, more dependencies and configurations will be needed.

3 Additional Knowledge: About transformers

1 What is transformers

In the field of Natural Language Processing (NLP),Transformers has become one of the most core model architectures. Originally proposed by Vaswani et al. in 2017, the core idea is throughSelf-Attention mechanism, allowing the model to understand the associations between different positions in the text without relying on traditional recurrent or convolutional structures. This mechanism enables the model to efficiently capture contextual information in long texts, making it suitable for various language understanding and generation tasks.

In the Hugging Face ecosystem, transformers refers not only to this model architecture, but also to an open-source Python library. The purpose of this library is:

Unified interface to load models:Whether it is GPT, BERT, RoBERTa, T5, or Mistral, you can use the same library to load, infer, and train.

Conveniently call various pre-trained models:There are thousands of open-source models on the Hugging Face Hub, covering tasks such as text generation, text vectorization (embedding), classification, question answering, and translation. The model will be automatically downloaded and cached locally upon the first call, and can then be run directly in Python scripts.

Flexible running locally or in the cloud:You can infer directly on local CPU/GPU, or combine with cloud services or inference platforms to deploy large models. This means that even without Ollama or OpenAI API, you can still run a complete RAG Demo using transformers.

In short, transformers is the core engine of the RAG Demo: It is responsible for converting text into vectors (Embedding) and can also be used to generate answers (LLM), covering almost all functions that might be used in the entire NLP workflow. In this article, I mainly use it to handle Embedding, but its capabilities go far beyond this, which also paves the way for subsequent expansion to more complex local RAG or LangChain/LlamaIndex.

2 Available Model Types on Hugging Face

In Hugging Face's transformers ecosystem, you can directly use multiple types of models. Each type has a specific purpose, and the most common types of models in RAG or general NLP scenarios are summarized below:

Model Type Typical Use Example Model Description
Embedding (Vectorization) Text vectorization, used for similarity retrieval and RAG vector databases sentence-transformers/all-MiniLM-L6-v2, all-mpnet-base-v2 Mainly responsible for converting text or sentences into fixed-dimension vectors, facilitating vector retrieval by libraries like FAISS
LLM (Large Language Model) Text generation, dialogue, summarization gpt2, mistralai/Mistral-7B-v0.1, LLaMA Can generate text locally, used for the “generate answer” step of RAG, or can be used independently for Q&A or creation
Text classification Sentiment analysis, topic classification, spam detection, etc. distilbert-base-uncased-finetuned-sst-2-english Outputs class probabilities, used to determine text attributes or labels
Question Answering (Extractive QA) Extracts specific answers from paragraphs deepset/roberta-base-squad2 Given a question and context, the model returns a text segment as the answer
Translation Translation between languages Helsinki-NLP/opus-mt-en-zh Supports translation for multiple language pairs, and can also be used in scenarios like cross-lingual retrieval
Multimodal Speech recognition, image + text, audio classification facebook/wav2vec2-base-960h, some CLIP models Requires extra dependencies (such as torchaudio, datasets), can process non-text data

💡 Supplementary Notes

  • Core of the RAG Demo: In the minimized RAG Demo of this article, we mainly use Embedding for vectorization, while optionally choosing a small LLM for generation; other types of models can serve as bases for subsequent extensions or other NLP projects.
  • Can run without Ollama: The models shown in this table can all be loaded and inferred locally using transformers, without relying on Ollama or any online platform.

3 Model Loading Methods: Hub vs. Local

When using transformers or other Hugging Face models, there are mainly two ways to load the model: directly from Hugging Face Hub download, or use those already saved to local disk . Note that these two methods may load the same model, just with different sources and loading methods.

Hub models are hosted on the official Hugging Face servers and can be downloaded directly via from_pretrained().Internet connection is required for the first load, and once downloaded, the model will be automatically cached locally, and subsequent calls usually no longer depend on the network. The advantage of Hub models is that they are very convenient to obtain and update, making them suitable for quick experimentation, learning, or testing new models without manual file management; the downside is thatthey rely on the network during initial download and updates, and if the Hugging Face service is temporarily unavailable, you may not be able to pull or upgrade models.

Local models refer to model files that have already been saved on the disk, which can be loaded by pointing to the local path. The benefit of using local models is that they are completely offline and highly controllable, making them more stable and reliable for formal deployment. The disadvantage is the need to manually manage model files; if you want to upgrade, you must manually replace or re-download them.

Overall, the difference between Hub models and local models mainly lies in convenience and controllability: Hub models are convenient and fast, suitable for learning and rapid iteration; local models are stable and controllable, suitable for formal deployment or completely offline scenarios. Understanding the pros and cons of these two loading methods helps in choosing the most appropriate strategy at different stages, and also lays the foundation for implementing the subsequent RAG workflow.

4 Recommendations for Formal Deployment

In production environments, if the goal is a stable, high-performance blog chatbot, choosing the model loading method and inference platform becomes particularly critical. For devices with sufficient hardware (such as the M4 Pro Mac mini),the Ollama method is usually the optimal choice. Ollama is deeply optimized for Apple M-series GPUs, fully utilizing GPU computing power, supporting low-precision computation (fp16/bf16) and efficient memory layouts, so even medium-to-large models can run efficiently locally. Meanwhile, Ollama also optimizes concurrent performance, making it more suitable for formal deployment scenarios.

On the other hand, using the HF transformers local methodis also completely feasible. It relies on PyTorch's Metal API for acceleration on Mac GPUs, which can significantly speed up inference for small models, but large models may be limited by GPU memory, requiring batched inference or hybrid CPU-GPU execution. The advantage of the HF method lies in its complete controllability and transparency, making it suitable for education, testing, or low-concurrency scenarios, allowing developers to deeply understand the internal logic of each step in the RAG workflow.

Overall, for the Demo or learning phase, you can simply use the local HF model to quickly run through the complete workflow of “document → embedding → retrieval → LLM answer generation” and easily understand how each step works. On the other hand, in production deployment or high-performance scenarios, if hardware permits, the Ollama approach is not only hassle-free but also offers the best performance, making it the more recommended solution; the local HF approach is still usable, but has limited performance and concurrency support for large models.

4 RAG Core Module Analysis and Example Code

1 The Five Steps and Three Major Roles of the Minimal RAG Closed Loop

In the introduction, I mentioned the five-step process of RAG: “chunking → vectorization → vector storage → retrieval → answer generation”. These five steps are specific operational steps, but from a higher level, the entire system can be abstracted into three core roles:text chunks, vector database, LLM。

Among them, “text chunk” does not refer to the original entire article, but to the smallest semantic unit obtained after chunking and vectorizing the original document—that is, the form in which knowledge is actually represented and retrieved in the system. In other words, the original document becomes text chunks through “chunking” and “vectorization”, and text chunks are the objects that the system needs to store and retrieve. The vector database is responsible for saving the vectors of these text chunks and performing similarity retrieval; the LLM is responsible for generating the final answer based on the retrieved context.

Therefore, the five steps are the pipeline at the “operational level”, while the three major roles are the participants at the “structural level”. After understanding this hierarchical relationship, the subsequent work becomes very clear: first, process the original document into an appropriate granularity oftext chunks(this is the first core), vectorize the text chunks and store them in the vector database (the second core), and then use the LLM combined with the retrieved results to generate answers (the third core). To make it easier for readers to follow, I will add category labels after the subsequent section titles (for example, 2 Preparing Input Documents (Text Chunk Category)、3 Document Chunking (Text Chunk Category) etc.), so that the content of each section and its corresponding role will be clear at a glance.

After understanding these three core roles, the next step is to prepare the first core:Text Chunks (Text Chunk Category), which means organizing the content you want the RAG system to “remember”, and turning it into units that the system can use through chunking and vectorization. In the following chapters, I will guide you step-by-step to build this minimum closed loop.

2 Text Chunks

2.1 Preparing Raw Documents

In the RAG workflow,text chunksis the smallest unit after knowledge enters the system. But before chunking, we first need to prepare an original document as the starting point for subsequent processing.

For the convenience of demonstration, I will directly select an existing .md file as the original document. There are two reasons for choosing a Markdown file: on one hand, it is a common knowledge management format, and many people use .md to write blogs and take notes; on the other hand, it is essentially still a plain text file, which is very simple to read and process, involving no complex format parsing.

It should be noted that there is no rigid requirement for the format of the original document; the .md file is just an example, and it could also be a .txt file, or even plain text scraped from web pages. The essence of RAG is processing text—as long as the content can be converted into a string, it can enter the subsequent workflow.

In this current minimal demo, we only use the simplest plain text files (such as .md, .txt) because Python can read them directly without relying on additional libraries. If you need to process complex formats like .docx, .pdf, etc., you will need to use additional parsing tools (such as python-docx, pymupdf, etc.), which are beyond the scope of this chapter.

Suppose we have a sample.md file in the current directory, the content of which might be an article or a note.In the following chapters, we will first chunk this document into text chunks, then vectorize them and store them in a vector database, gradually building a minimum closed-loop RAG system.

2.2 Document Splitting

Once we have the original document, the first step is chunkingThe reason is simple: most documents are often very long, even containing thousands of words. If they are directly fed into the vectorization model, it is not only computationally inefficient but also leads to overly vague semantic representation, making it difficult to hit precise segments during retrieval.

So-called chunking is splitting a long document into several smaller text chunksEach chunk should keep its semantic integrity as much as possible without being too large; otherwise, it may still exceed the model's processing limit during subsequent embedding. Usually, a chunk length of a few hundred to a thousand charactersis more appropriate.

In the world of RAG, chunking is a crucial step: if chunks are too small, the semantics may be fragmented; if they are too long, it is unfavorable for retrieval and matching. Here, we do not pursue complex chunking algorithms, but only make a minimal demo. The simplest way is to split by paragraph, such as chunking once whenever a newline character is encountered. In this way, the logical structure can be basically preserved, while the document is cut into small pieces.

In the subsequent code implementation, we will write a short function to complete this task: read sample.md, split it by paragraph, and return a list of text chunks. This list will be the input for the subsequent vectorization.

In other words,Chunking is the “entrance of the entrance”, which determines the granularity at which knowledge enters the RAG system. Only with reasonable chunking can the results of retrieval and answer generation be satisfying.

2.3 Vectorization

In the core role of “text chunks”, chunking is only the first step, which allows long documents to be broken down into smaller fragments. But text chunks alone are not enough; machines cannot directly understand text. To give them “semantic awareness”, we need to further convert these text chunks into vector representations。

So-called vectorization is the process of encoding natural language into a sequence of numbers—these numbers correspond to the semantic coordinates of the model in a high-dimensional space, which can be used to measure the similarity between different text chunks.

To give a simple example: when humans see “apple” and “banana”, they know they both belong to fruits; while “apple” and “laptop” are completely different literally, we can still understand that “apple” refers to Apple Inc. in certain contexts. The role of vectorization is to let machines also perceive this “semantic distance” in the digital world. Two text chunks with close meanings will have very close vectors; conversely, the distance will be pulled apart. (For a detailed introduction to vectors, please refer to the article: ...)

In practice, we can leverage off-the-shelf embedding models provided by Hugging Face, such as sentence-transformers/all-MiniLM-L6-v2The advantage of this model is that it is compact, fast, and has sufficient accuracy, making it very suitable for demos. Through the transformers library, we can easily convert text chunks into vectors.

At this point, the work of the ”text chunk” role comes to an end. We already have a set of vectorized “semantic numbers” in hand, but they are still just data scattered in memory. Next, we need to introduce the second core role—vector database, to organize these vectors for easy retrieval and answer generation.

3 Vector Database

3.1 The Role of the Vector Database

In the previous section, we already converted text chunks into vectors. At first glance, these vectors are just a bunch of high-dimensional numerical arrays, which are actually of little use when existing in isolation. The real value lies in the fact that we can put them into a specialized storage and retrieval tool, which is the vector database.

Why do we need a vector database? Imagine: if we have a few hundred text chunks, directly using Python to traverse the array and calculate ”cosine similarity” is still acceptable; but if the amount of data becomes hundreds of thousands or even millions of chunks, comparing them one by one will become extremely slow and almost unusable. The role of a vector database is to provide efficient similarity search, allowing us to quickly find the few chunks closest to the query in a massive collection of vectors.


The term ”cosine similarity” was mentioned in the passage above. What is the specific meaning of this name?

Simply put, after each text chunk is vectorized, it becomes a point in a high-dimensional space. If we think of a vector as an arrow extending from the origin, then “cosine similarity” is comparing the angle between the two arrows: the smaller the angle (the closer the direction), the more semantically similar the two text chunks are; the larger the angle (the greater the difference in direction), the larger the semantic difference between them.

Because cosine similarity only cares about direction and not the length of the vector, it is particularly suitable for measuring semantic relevance. In other words, a vector library is like a “semantic index”. Unlike traditional databases that use keywords for retrieval, it finds relevant content through “semantic distance”. For example, if a user asks: “When will the new Apple product be released?”—the vector library will look for semantically related chunks like “Apple Event” instead of rigidly matching the word “Apple”.


In practical applications, there are many “vector libraries” to choose from, but it is worth noting that the “vector libraries” commonly referred to in the industry usually include two categories:Vector search libraries (vector search library) 和 Vector databases (vector database)The former focuses more on efficient similarity retrieval, while the latter adds data management and scalability on top of that:

  • Vector search libraries—These are the underlying retrieval algorithms and index implementations, with typical representatives being FAISS, Annoy, hnswlibThey provide efficient nearest neighbor search (K-NN), multiple index structures, and compression strategies, making them suitable for embedding into local applications or serving as core components for vector retrieval. The advantages are lightweight, good performance, and zero O&M cost (single machine), while the disadvantages are that they usually do not include complete database functions (such as complex metadata filtering, permission management, distributed scaling, etc.). If you use FAISS for retrieval but need to save the document ID, original text, or other metadata corresponding to each vector, you usually need to use an external small database (such as SQLite, LevelDB, or a simple JSON/CSV) to store this meta-information.
  • Vector databases—This is a productized wrapper on top of search libraries, which contains both high-performance retrieval capabilities and provides database-level features: metadata indexing/filtering, persistence, sharding/replication, online index management, query interfaces, cloud-hosted services, etc. Typical representatives include Milvus, Qdrant, Weaviate, PineconeThe advantages are full-featured, easy for production-grade deployment and scaling; the disadvantages are heavyweight, higher O&M or costs (especially cloud services). Among them, Milvus, Qdrant are more commonly used as self-hosted/open-source deployments;Pinecone, Weaviate (managed version) are cloud services, hassle-free but paid.

Example comparison (when to choose which):

  • For example, if you are doing introductory experiments or small-scale local RAG in your own environment and want “zero O&M, rapid validation”,FAISS (vector search library) is the preferred choice, as data management and scalability are not required.
  • If you want to build online services, need metadata filtering (such as filtering candidate segments by date, author, category), multi-node scaling, or have a stable SLA, it is recommended to consider Vector databases(Milvus, Qdrant, Pinecone, etc.).

In my minimalist RAG Demo, I will choose FAISS (i.e., the previously installed faiss-cpu): It has few dependencies, simple configuration, and can intuitively demonstrate semantic retrieval effects, making it highly suitable for local experiments on a Mac mini. If you need to upgrade the Demo to a production-grade system in the future, you can then consider replacing or connecting FAISS to a vector database to obtain more management and scaling capabilities.

3.2 Building Vector Index

After the text chunks are converted into vector representations, I need to organize them into a data structure that can be quickly retrieved. This is the role of vector index , which can be understood as the core component in a vector library—it is responsible for efficiently storing vectors and supporting similarity queries.


When building a vector index, we can usually choose in-memory index, on-disk index, or a combination of both. Below areCommon types and storage methods of vector indexes:

Use Cases Storage Type Common Index Features & Applicable Scenarios
Small-scale experiments / debugging In-memory index IndexFlatIP / IndexFlatL2 Exact search, simple structure, fast query; but memory consumption rises rapidly with data volume, suitable for Demo and validation
Million-scale retrieval In-memory + approximate index IVFFlat / IVFPQ Approximate search using clustering and quantization, sacrificing some accuracy for speed and storage, suitable for large-scale retrieval
Online high-performance services In-memory graph index HNSW Graph-based approximate search, low latency, high recall rate, suitable for applications with high real-time requirements
Massive data persistence On-disk index or hybrid index DiskANN / Faiss on-disk IVF Supports persistence, remains available after restart, suitable for TB-level vector storage and offline/near-line queries

For a local RAG Demo, if the ultimate goal is to write vectors uniformly into a vector database, then directly using an in-memory index during index construction and debugging is absolutely fine.


Next, I will use a Python script to demonstrate a minimal Demo, using FAISS(Facebook's open-source vector search library, which is highly performant and easy to run locally) to implement vector indexing, demonstrating how to add vectors to the index and perform similarity queries. For ease of understanding, I will first break down the key steps of the script into steps 1, 2, and 3; in practice, you can directly run the complete script in step 4.

Initialize the vector index

import faiss
import numpy as np

dimension = 384  # 与文本向量维度一致
index = faiss.IndexFlatIP(dimension)  # 使用内积,可用于余弦相似度

Here, IndexFlatIP is used because cosine similarity can be converted into an inner product calculation after vector normalization, allowing the vector of each text chunk to be added directly to the index.

Add vectors to the index

# 假设已有向量列表 vec1, vec2, vec3
vectors = np.array([vec1, vec2, vec3], dtype='float32')
index.add(vectors)  # 将向量加入索引
print(f"向量索引中共有 {index.ntotal} 个向量")

Once added, the vector index is built and can be immediately used for similarity retrieval.

Persistence (Optional)

# 保存索引到磁盘
faiss.write_index(index, "vector_index.faiss")

# 下次加载
index = faiss.read_index("vector_index.faiss")

This way, even if the program ends, the existing vector index can be used directly upon the next startup, achieving persistence.

At this point, we have completed the construction of a minimal vector database,which provides a reliable storage foundation for subsequent text retrieval and answer generationIn the next section, we will discuss how to perform incremental management, saving, and loading of the vector database to cope with the continuously growing text chunks in practical use.

3.3 Vector Database Management

Once the vector database is built, it is not a matter of “set it and forget it.” In practical use, the vector database needs to be dynamically managed to ensure retrieval efficiency and data integrity. Management mainly includes incremental addition, updating, saving, and loading in four aspects:

  1. Incremental Addition

As new text chunks continuously generate new vectors, the vector database should support adding them at any time. FAISS supports directly calling index.add() to incrementally add vectors:

new_vectors = np.array([vec_new1, vec_new2], dtype='float32')
index.add(new_vectors)
print(f"向量库总数更新为 {index.ntotal}")

This approach does not require rebuilding the index, making it highly suitable for scenarios in RAG where knowledge is dynamically expanded.

  1. Vector Updating

If a text chunk is modified, the corresponding vector needs to be deleted first and then re-added. FAISS native indexes (such as IndexFlatIP) do not support single-item deletion, but this can be achieved using IndexIDMap combined with custom IDs:

# 创建带 ID 的索引
index_id = faiss.IndexIDMap(index)
index_id.add_with_ids(vectors, ids)
# 更新向量时,先 remove 再 add
  1. Saving and Loading

To prevent data loss due to program exit, the vector database should be saved regularly:

faiss.write_index(index_id, "vector_index.faiss")
# 下次加载
index_id = faiss.read_index("vector_index.faiss")

For large indexes, segmented saving or incremental saving strategies can be combined to improve efficiency.

  1. Index Optimization (Optional)

For large-scale vector databases, FAISS's IVF, PQ, and other compression or clustering indexes can be used to improve retrieval speed while saving memory:

nlist = 100  # 聚类数
quantizer = faiss.IndexFlatL2(dimension)
index_ivf = faiss.IndexIVFFlat(quantizer, dimension, nlist, faiss.METRIC_L2)
index_ivf.train(vectors)
index_ivf.add(vectors)

Although complex, it is highly necessary when the number of vectors grows to hundreds of thousands or millions. Through these management methods, the vector database can continuously and stably serve RAG retrieval, and even if text chunks continuously increase, it will not affect retrieval efficiency and accuracy.

At this point, the prototype of the vector database has been built. We can not only store the vectors of text chunks in an orderly manner but also retrieve them quickly when needed, and make them still available after program restarts through persistence.With such a stable “semantic repository,” the infrastructure of RAG is gradually becoming complete.Next, we enter the final part of Chapter 4—how to use the vector database to retrieve information and ultimately generate the answers we want.

4 Large Language Model

4.1 The Role of Large Language Models in RAG

In the previous two sections, we completed the vectorization of “text chunks” and organized these vectors into the “vector database.” In this way, after retrieving the user's question, a set of semantically related text fragments can be found. Next, it is the turn of Large Language Models (LLMs) to take the stage.

In the RAG architecture, the main responsibilities of the LLM can be summarized into three points:

  1. Understanding user questions

The LLM must first semantically understand the questions input by the user and identify the information needs behind them.

  1. Combining retrieval results

Relying solely on the LLM itself, its knowledge may be outdated or limited. Therefore, we need to provide the “relevant text chunks” retrieved from the vector database as supplementary information for the LLM's reference. In this way, the model's answers can be based on the latest, domain-specific data, rather than just relying on training corpora.

  1. Generating the final answer

After obtaining the question and supplementary materials, the LLM is responsible for combining the two to generate a natural language answer. This step must ensure both factual correctness and fluent language, reflecting the value of RAG.

The role of the LLM in RAG can be compared to a “final interpreter”: the vector database provides “reference materials,” while the LLM decides how to organize these materials to ultimately provide the most useful answer to the user.

To enable the large model to smoothly fulfill these responsibilities, we must design a reasonable input structure (prompt structure), which is how to combine the “user question + retrieved text chunks” and input them into the large model.

4.2 Input Structure

In the previous subsection, we defined the three responsibilities of the large model in the RAG architecture: understanding the question, combining the retrieval results, and generating answers. To enable the large model to effectively complete these tasks, the key lies in the design of input structure (prompt structure) .

In other words, the kind of answer the large model can output largely depends on how we concatenate the user question and the retrieved text chunks. If the input structure is messy, the model may ignore the retrieval results; if the input is too verbose, the model may lose the focus.

A common input structure generally includes three parts:

  1. Instruction part (Instruction)

Clearly tell the large model what its task is, for example: “Please answer the user's question based on the following materials, and do not make things up when answering.” This part acts as setting the tone to prevent the model from going off-topic or hallucinating.

  1. Context part (Context)

This places several relevant text chunks retrieved from the vector database. They are the “factual basis” for the large model to generate answers, acting as temporary external knowledge.

  1. Question part (Question)

Finally is the user's original question. Placing the question near the end helps the model focus on the question itself after understanding the context.

A typical input structure can look like this:

你是一名智能助手,请严格根据“上下文”回答用户的问题。如果答案无法在上下文中找到,请明确回答“我不知道”,不要编造。

【上下文】
{text_chunks}

【问题】
{user_question}

The benefit of this structure is: it has clear task instructions to prevent the model from improvising arbitrarily; the context and question are clearly partitioned, making it easy for the model to parse; it can maintain the relevance of the generated content while reducing the risk of hallucination.

Of course, in actual projects, the input structure can be continuously optimized according to requirements, such as adding “answer format requirements” or controlling “answer length”. But in any case, the core idea remains consistent:Enable the large model to combine the retrieved materials and the user question to generate answers under a clear task framework。

4.3 Retrieval-Augmented Invocation Workflow

Having understood the responsibilities of the large model and the basic form of the input structure, you can now connect them with the vector database to form a complete retrieval-augmented calling workflow. The core of this workflow is:User asks a question → Retrieve relevant documents → Construct input → Large model generates the answer。

The entire process can be divided into four steps:

User asks a question

The user inputs a natural language question, for example: “What is the difference between a vector database and a vector search library?”

Vector retrieval

The system first converts the question into a vector, and then retrieves the most similar text chunks from the vector database. These text chunks are “candidate knowledge” and will be provided as context to the large model.

Construct the input structure

According to the rules of the previous section, concatenate the retrieved text chunks (Context) and the user question (Question) into a unified input, and add task instructions (Instruction).

Large model generates the answer

Finally, hand over the concatenated input to the large model and let it generate an answer based on the context. If the context does not cover the information required for the question, the model will answer “I don't know” according to the instructions.

A simplified pseudo-code workflow looks something like this:

# 1. 用户输入
question = "向量数据库和向量搜索库有什么区别?"

# 2. 向量检索
q_vector = embed(question)  # 将问题向量化
D, I = index.search(q_vector, k=3)  # 在向量库中检索 top-3
retrieved_chunks = [chunks[i] for i in I[0]]

# 3. 构造输入
prompt = f"""
你是一名智能助手,请根据以下“上下文”回答问题。
如果答案无法在上下文中找到,请回答“我不知道”。

【上下文】
{retrieved_chunks}

【问题】
{question}
"""

# 4. 大模型生成答案
answer = llm.generate(prompt)
print(answer)

Thus, a complete RAG calling closed-loopis established:

  • The vector database provides external knowledge;
  • The input structure ensures that knowledge is correctly transmitted;
  • The large model is responsible for understanding and generating natural language answers.

This is also why we say that RAG is not “letting the large model know everything”, but “letting the large model learn to use external materials”.

4.4 Additional Knowledge: Selection of Large Language Models

In the previous subsections, we have clarified the responsibilities of the large model in RAG, the input structure, and the retrieval-augmented calling workflow. The next step is to consider which specific large model to use(for the sake of brevity in this Demo, the LLM model provided by Hugging Face is used for demonstration; if you pursue performance or want local deployment, you can also run models like LLaMA through the Ollama platform). This choice will directly affect the generation quality and speed of the model, as well as hardware resource requirements.

Model selection principles

When choosing a large model, the following factors generally need to be considered:

Generation capability and accuracy

The larger the model, the stronger its ability to understand questions and generate high-quality answers usually is, but it also consumes more resources.

Hardware resource limitations

Including GPU/CPU memory, thread count, whether quantization is supported, etc.

延迟与响应速度

实时交互场景对生成速度有一定要求,模型过大可能导致响应延迟。

可用性与兼容性

是否有可量化版本(GGUF、GGML、bitsandbytes 等)、是否易于本地部署或通过 Hugging Face 加载。

2、常见可选模型对比

Model 参数量 Advantages 硬件建议
GPT-2 / GPT-Neo 125M–1.3B 0.1–1.3B 小巧,推理快,容易部署 Mac/PC CPU 或小型 GPU 均可运行
Mistral 7B 7B 高性能、推理速度快 16GB+ GPU 或 24GB Mac 内存
LLaMA 3 8B(量化 GGUF) 8B 兼顾性能与占用,适合本地部署 24GB RAM 的 Mac M 系列可运行
LLaMA 3 13B 13B 生成能力更强,适合复杂问题 40GB+ GPU 或大型服务器
LLaMA 3 70B 70B 极高生成能力,但资源消耗大 多 GPU 或大型云服务器,本地 Mac 不适用

注:模型越大,所需的显存和内存越多。在本地部署时,选择量化版本(如 Q4_K_M GGUF)可以大幅降低资源占用,同时保持合理生成效果。

3、选择策略

  • 个人笔记本或本地 Mac

推荐 LLaMA 3 8B(量化版 GGUF),兼顾效果与速度。

  • 中型 GPU 服务器或云环境

可以选择 Mistral 7B 或 LLaMA 3 13B,获得更强生成能力。

  • 高端多 GPU 云环境

可考虑 LLaMA 3 70B 或其他超大模型,追求极致生成效果。

4、实践建议

  • 根据你的硬件资源量身选择模型,避免因内存不足导致运行失败。
  • 对于 RAG Demo 这种交互式演示,模型不必过大,保持响应速度和可运行性即可。
  • 如果未来需要处理更大规模知识或复杂问题,再考虑升级到更大模型。

5 Script Structure and Data Flow

1 Script Structure Explanation

在前面的章节里,我们分别从文档块、向量库和大模型三个角色的角度,逐步拆解了最小 RAG Demo 的关键环节。为了把这些分散的步骤真正串联起来,我们需要把它们整理成几个独立的 Python 脚本。这样做的好处是:逻辑清晰、职责明确,也方便日后扩展或替换某一部分。

在本次 Demo 中,我们会使用三个脚本:

1. document_process.py

  • 负责:文档切分与向量化(对应第 4.2 章的“文本块”)。
  • 输入:原始文档(txt/markdown 等)。
  • 输出:向量化后的文本块列表,存放在内存或中间文件中。

2. vector_store.py

  • 负责:构建向量索引与管理(对应第 4.3 章的“向量库”)。
  • 输入:document_process.py 生成的向量。
  • 输出:完成索引构建的向量库(支持检索,并可选择持久化到本地文件)。

3. rag_pipeline.py

  • 负责:组织大模型调用,结合检索结果生成答案(对应第 4.4 章的“大模型”)。
  • 输入:用户的查询问题。
  • 输出:RAG 的最终生成结果。

这三个脚本既能独立运行(方便测试),又能通过数据传递形成一个完整的流水线。换句话说:它们和“三大核心角色”是完全对应的,只是实现上稍微拆分了任务。

2 document_process.py

"""
document_process.py

功能:完成“文档块”的处理工作(优化版)
包括:
1. 文档切分(对应第 4.2 文档块 - 文档切分)
2. 文本向量化(对应第 4.2 文档块 - 向量化)
3. 输出过程文件:processed_docs.json + embeddings.npy
"""

import os
import json
import numpy as np
from sentence_transformers import SentenceTransformer

# ======================
# 1. 文档切分
# ======================
def load_and_split_md(file_path):
    """加载 Markdown 文件,并按段落切分"""
    blocks = []
    with open(file_path, "r", encoding="utf-8") as f:
        paragraph = []
        for line in f:
            line = line.strip()
            if line == "":
                if paragraph:
                    blocks.append(" ".join(paragraph))
                    paragraph = []
            else:
                paragraph.append(line)
        if paragraph:
            blocks.append(" ".join(paragraph))
    return blocks

def load_and_split_dir(dir_path):
    """遍历目录及子目录下所有 .md 文件,并按段落切分"""
    all_blocks = []
    for root, dirs, files in os.walk(dir_path):
        for file in files:
            if file.lower().endswith(".md"):
                file_path = os.path.join(root, file)
                blocks = load_and_split_md(file_path)
                all_blocks.extend(blocks)
    return all_blocks

# ======================
# 2. 文本向量化
# ======================
def embed_texts(texts, model_name="sentence-transformers/all-MiniLM-L6-v2"):
    """使用预训练模型将文本块转为向量"""
    model = SentenceTransformer(model_name)
    embeddings = model.encode(texts, convert_to_numpy=True, normalize_embeddings=True)
    return embeddings

# ======================
# 主流程
# ======================
if __name__ == "__main__":
    dir_path = "document"
    print(f"正在处理目录: {dir_path}")

    # Step 1: 遍历目录,生成文本块
    text_blocks = load_and_split_dir(dir_path)
    print(f"共生成 {len(text_blocks)} 个文本块。")

    # 保存文本块 JSON
    with open("processed_docs.json", "w", encoding="utf-8") as f:
        json.dump(text_blocks, f, ensure_ascii=False, indent=2)
    print("文本块已保存至 processed_docs.json")

    # Step 2: 文本向量化
    vectors = embed_texts(text_blocks)
    np.save("embeddings.npy", vectors)
    print("向量已保存至 embeddings.npy")
    print("向量化完成,向量矩阵维度:", vectors.shape)

document_process.py 脚本说明:

这个脚本做了两件事:

1. 文档切分(对应 4.2.1 文档块 – 文档切分):遍历 document/ 目录下的 Markdown 文件,将每个文档按段落拆成小块,并生成 processed_docs.json 文件,记录切分后的所有文本块。

2. 文本向量化(对应 4.2.2 文档块 – 向量化):调用 sentence-transformers 模型,把文本块转成向量,并将向量以 NumPy 数组形式保存在 embeddings.npy 文件中,供下一步构建向量索引使用。

生成的中间文件:

  • processed_docs.json:切分后的文档片段列表,用于构建向量索引。
  • embeddings.npy:文本块对应的向量矩阵,用于构建向量索引。

运行方式:

  • 保存为 document_process.py
  • 在终端运行:
python3.11 document_process.py

3 vector_index.py

"""
vector_index.py

功能:构建向量索引并管理
包括:
1. 加载文本块和向量文件
2. 构建 FAISS 向量索引
3. 保存索引及映射关系
"""

import faiss
import numpy as np
import pickle
import json

# ======================
# 参数设置
# ======================
dimension = 384
index_file = "vector_index.faiss"
mapping_file = "vector_index.pkl"
text_blocks_file = "processed_docs.json"
vectors_file = "embeddings.npy"
top_k_demo = 3  # 演示检索前 k 个结果

# ======================
# Step 1: 加载文本块与向量
# ======================
with open(text_blocks_file, "r", encoding="utf-8") as f:
    text_blocks = json.load(f)
vectors = np.load(vectors_file)
print(f"加载 {len(text_blocks)} 个文本块和向量矩阵,维度 {vectors.shape}")

# ======================
# Step 2: 构建 FAISS 索引
# ======================
index = faiss.IndexFlatIP(dimension)  # 内积作为余弦相似度
index.add(vectors)
print(f"向量索引构建完成,总向量数: {index.ntotal}")

# ======================
# Step 3: 保存索引和映射关系
# ======================
faiss.write_index(index, index_file)
print(f"索引已保存至 {index_file}")

with open(mapping_file, "wb") as f:
    pickle.dump(text_blocks, f)
print(f"向量映射关系已保存至 {mapping_file}")

# ======================
# Step 4: 简单检索演示
# ======================
query_vec = vectors[0].reshape(1, -1)
distances, indices = index.search(query_vec, top_k_demo)

print("\n检索演示:")
print(f"查询文本块: {text_blocks[0]}\n")
for rank, idx in enumerate(indices[0], 1):
    print(f"Rank {rank}: 相似度 {distances[0][rank-1]:.4f}")
    print(f"文本块内容: {text_blocks[idx]}\n")

vector_index.py 脚本说明:

这个脚本做了两件事:

1. 构建向量索引(对应 4.3 向量库 – 构建索引):

  • 从 processed_docs.json 中加载文本块
  • 从 embeddings.npy 中加载文本块向量
  • 使用 FAISS 将向量加入向量库
  • 建立向量与文本块的映射关系

2. 向量库持久化(对应 4.3 向量库 – 持久化管理):

  • 将构建好的向量索引写入磁盘,生成 vector_index.faiss
  • 将向量与文本块的映射关系保存为 vector_index.pkl,方便后续检索时对应文本块

生成的中间文件:

  • vector_index.faiss:向量索引文件,用于快速检索
  • vector_index.pkl:向量与文本块映射关系,用于将检索结果还原为原始文本块

运行方式:

  • 保存为 vector_index.py
  • 确保 document_process.py 已生成 processed_docs.json 和 embeddings.npy
  • 在终端运行:
python3.11 vector_index.py

4 rag_query.py

"""
rag_query.py

功能:完成“大模型”在 RAG 中的调用
包括:
1. 加载向量索引与映射关系
2. 接收用户查询
3. 基于相似度检索相关文本块
4. 调用大模型生成答案,并展示参考片段(控制台美化输出)
"""

import faiss
import pickle
from sentence_transformers import SentenceTransformer
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

# 尝试导入 colorama,若不可用则回退为空字符串(兼容无 colorama 环境)
try:
    from colorama import Fore, Style, init
    init(autoreset=True)
except Exception:
    class _Dummy:
        def __getattr__(self, name):
            return ""
    Fore = _Dummy()
    Style = _Dummy()

# ======================
# 参数设置
# ======================
index_file = "vector_index.faiss"
mapping_file = "vector_index.pkl"
embedding_model_name = "sentence-transformers/all-MiniLM-L6-v2"

# 中文优化的公开模型
lm_model_name = "Jackrong/llama-3.2-3B-Chinese-Elite-v2"
top_k = 3

# ======================
# Step 1: 加载索引与映射
# ======================
index = faiss.read_index(index_file)
with open(mapping_file, "rb") as f:
    text_blocks = pickle.load(f)
print(f"{Fore.CYAN}向量索引加载完成,总向量数: {index.ntotal}{Style.RESET_ALL}")
print(f"{Fore.CYAN}映射文本块加载完成,总数: {len(text_blocks)}{Style.RESET_ALL}")

# ======================
# Step 2: 初始化查询向量模型
# ======================
embed_model = SentenceTransformer(embedding_model_name)

# ======================
# Step 3: 初始化 LLM(适配 macOS / CPU)
# ======================
tokenizer = AutoTokenizer.from_pretrained(lm_model_name)
lm_model = AutoModelForCausalLM.from_pretrained(
    lm_model_name,
    device_map="auto",
    torch_dtype=torch.float16
)

# ======================
# Step 4: 查询与检索
# ======================
def retrieve_relevant_blocks(query, top_k=3):
    query_vec = embed_model.encode([query], convert_to_numpy=True, normalize_embeddings=True)
    distances, indices = index.search(query_vec, top_k)
    results = [(text_blocks[i], float(distances[0][rank])) for rank, i in enumerate(indices[0])]
    return results

# ======================
# Step 5: 生成答案(优化 prompt)
# ======================
def generate_answer(query):
    retrieved = retrieve_relevant_blocks(query, top_k)
    context = "\n".join([block for block, _ in retrieved])

    prompt = f"""
你是一名智能助手。你的任务是根据给定的“上下文”回答用户问题。
请注意:
- 严格引用上下文中的关键句回答问题;
- 将多条信息合并成一句或两句,尽量精简,避免重复;
- 不要复述上下文或问题;
- 如果上下文中找不到答案,请输出“我不知道”;
- 答案以 <END> 结束。

上下文:
{context}

问题:
{query}

答案:<START>"""

    inputs = tokenizer(prompt, return_tensors="pt")
    inputs = {k: v.to(lm_model.device) for k, v in inputs.items()}

    output_ids = lm_model.generate(
        **inputs,
        max_new_tokens=200,
        do_sample=False,
        repetition_penalty=1.2,
        eos_token_id=tokenizer.eos_token_id
    )
    full_output = tokenizer.decode(output_ids[0], skip_special_tokens=True)
    answer = full_output.split("<START>")[-1].split("<END>")[0].strip()

    return answer, retrieved

# ======================
# Step 6: 交互式演示
# ======================
if __name__ == "__main__":
    while True:
        user_query = input(Fore.YELLOW + "请输入你的问题(回车退出):" + Style.RESET_ALL)
        if not user_query.strip():
            break
        answer, refs = generate_answer(user_query)

        print(f"\n{Fore.GREEN}RAG 生成答案:{Style.RESET_ALL}\n{Fore.GREEN}{answer}{Style.RESET_ALL}")
        print(f"\n{Fore.MAGENTA}参考片段:{Style.RESET_ALL}")
        for rank, (block, score) in enumerate(refs, 1):
            print(f"{Fore.YELLOW}[Rank {rank}] 相似度: {score:.4f}{Style.RESET_ALL}")
            print(f"{Fore.LIGHTBLACK_EX}{block}{Style.RESET_ALL}")
            print(f"{Fore.CYAN}{'-' * 40}{Style.RESET_ALL}")
        print(f"{Fore.CYAN}{'=' * 80}{Style.RESET_ALL}")

这个 rag_query.py 脚本 实现了一个简易的 RAG流程,主要功能包括:

1、加载向量索引与映射关系

  • 从 vector_index.faiss 加载已构建的向量索引
  • 从 vector_index.pkl 加载文本块映射关系
  • 恢复已处理的文档向量库与文本块对应关系

2、查询文本块

  • 接收用户输入问题
  • 使用 SentenceTransformer 将问题向量化
  • 在向量索引中检索与问题最相似的文本块(默认返回 top_k=3)

3、调用大模型生成答案

  • 将检索到的文本块拼接成上下文
  • 构造 Prompt:只允许基于上下文回答问题
  • 调用大模型生成最终答案,并返回生成答案时参考的片段
  • 若上下文中无法找到答案,则返回 “我不知道”

4、交互式查询演示(带控制台美化)

  • 支持循环输入问题
  • 输出结果时进行了颜色区分(如果安装了colorama依赖):绿色 → 生成的答案;紫色/黄色 → 参考片段标题与相似度分数;灰色 → 参考文本块正文;青色 → 分隔符。
  • 让查询结果更加直观

5、Prompt 设计原则

  • 回答必须严格基于“上下文”文本块
  • 如果上下文中没有答案,返回 “我不知道”
  • 避免模型凭空生成或扩展内容

依赖的中间文件

  • vector_index.faiss:向量索引文件(由 vector_index.py 生成)
  • vector_index.pkl:向量映射关系(由 vector_index.py 生成)

运行方式:

python3.11 rag_query.py

注意事项与常见问题说明:

  1. 模型量化问题
    • 早期脚本中使用 load_in_8bit=True 或 load_in_4bit=True 进行量化加载, 在 Mac CPU / M 系列上可能会报错,因为需要依赖 bitsandbytes 或 accelerate。
    • 最新 Transformers 版本推荐使用 quantization_config 配置对象来代替旧参数。
    • 在本脚本中,已经去掉了 bitsandbytes 依赖,直接使用 CPU 加载模型,避免量化报错。
  2. 运行报错问题
    • 如果遇到 ImportError: CUDA not available 或 bitsandbytes 相关错误, 多半是因为旧量化方式需要 CUDA 支持。
    • 在本脚本中无需 GPU,直接用 CPU 执行,已适配 Mac M 系列 CPU。
    • 部分 transformers 版本可能提示某些参数无效(如 temperature、top_p),不影响正常运行,可忽略。
  3. 模型下载中断或下载速度慢
    • 大模型(尤其是 3B+ 模型)在 Hugging Face 下载时可能出现中断或网络问题, 会提示重试或下载失败。
    • 遇到此类问题,可以手动在浏览器或使用 git lfs 下载模型文件到本地,并在 from_pretrained() 中指定本地路径。示例:lm_model_name = “/path/to/local/model/directory”
  4. 建议
    • 对于 Mac CPU / M 系列用户,建议使用 1B 或 3B 的中文优化公开模型,避免 7B 或以上模型导致内存压力过大。
    • 若仅测试 RAG 流程,使用较小模型即可满足功能验证。

5 Summary

根据前面几个小节的内容,最终项目目录结构如下:

project_root/
├── document/                 # 原始文档目录(放 .md 文件),也可以嵌套子目录
│   ├── doc1.md
│   ├── doc2.md
│   └── …
├── processed_docs.json       # 文本块 JSON(由 document_process.py 生成)
├── embeddings.npy            # 文本向量矩阵(由 document_process.py 生成)
├── vector_index.faiss        # FAISS 向量索引文件(由 vector_index.py 生成)
├── vector_index.pkl          # 向量映射关系(由 vector_index.py 生成)
├── document_process.py       # 文档切分与向量化脚本
├── vector_index.py           # 构建向量索引脚本
└── rag_query.py              # 交互式问答脚本

实际数据流过程如下:

1. **文档切分与向量化(document_process.py)**
   - 输入:`document/` 下的所有 `.md` 文件  
   - 输出: 
       - `processed_docs.json`(切分后的文本块)  
       - `embeddings.npy`(对应文本块的向量矩阵)

2. **构建向量索引(vector_index.py)**
   - 输入:`processed_docs.json` + `embeddings.npy`  
   - 输出: 
       - `vector_index.faiss`(FAISS 向量索引)  
       - `vector_index.pkl`(文本块与向量映射关系)

3. **交互式问答(rag_query.py)**
   - 输入:用户问题 + 向量索引文件  
   - 过程:检索相关文本块 → 结合 LLM 生成答案  
   - 输出:在命令行交互式显示最终答案

6 Practical Exercise — Getting RAG Up and Running

在前一章我们已经准备好了三个核心脚本(document_process.py、vector_index.py、rag_query.py),以及用于测试的 Markdown 文档。下面的步骤将演示如何实际运行整个 RAG Demo:

  1. 新建项目目录及文档目录
mkdir ~/Projects
mkdir ~/Projects/document
  1. 将准备好的3个核心脚本复制到项目目录中(也可直接在目录中新建)
cp /xx/document_process.py ~/Projects
cp /xx/vector_index.py ~/Projects
cp /xx/rag_query.py ~/Projects
  1. 在文档目录中放置测试文档

将若干 .md 文件放入 ~/Projects/document/ 文档目录中(可以包含子目录),这些文件就是后续问答的知识来源,本次实操我放入了3篇文章对应.md文件:

image.png

4. 切分与向量化

在终端依次运行:

python3.11 document_process.py

完成后终端输出如下:

image.png

再运行:

python3.11 vector_index.py

完成后终端输出如下:

image.png

2个脚本都运行完成后会,在项目根目录生成 processed_docs.json和embeddings.npy(由document_process.py脚本生成)、vector_index.faiss 和 vector_index.pkl(由vector_index.py脚本生成)。

  1. 交互式问答

执行:

python3.11 rag_query.py

image.png

进入交互界面后输入问题,即可得到基于本地文档检索的答案,以下是我针对不同文章内容准备的一些测试问题及相应的输出:

针对文章 1(博客与 AI 时代价值)

  • 为什么在 AI 时代写个人博客仍然有价值?

RAG的响应如下:

image.png

  • 什么是“可信知识锚点”?它的意义是什么?

RAG的响应如下:

image.png

针对文章 2(Cloudflare Tunnel)

  • 使用 Cloudflare Tunnel 建站可能存在哪些 SEO 隐患?

RAG响应如下:

image.png

  • 为什么非标准端口会影响搜索引擎收录?

RAG响应如下:

image.png

针对文章 3(WordPress 多活架构)

  • 个人博客为什么要考虑 WordPress 多活架构?

RAG响应如下:

image.png

  • ** WordPress 多活架构的关键技术点有哪些?**

RAG响应如下:

image.png

从测试结果来看,RAG demo的效果是让我满意的。


RAG 的强项是在“针对单个问题,结合相关片段”来给出回答,也就是做“点对点”的知识检索和问答。

如果尝试让它一次性跨越多篇文章、甚至跨主题去做聚合总结,效果往往不稳定——因为模型需要在多个上下文间做“逻辑缝合”,这超出了 RAG 的舒适区。

所以在实际应用中,建议把问题聚焦到某个知识点或某篇文章上,逐步获得答案,再由你自己来做整合。这样才能最大化 RAG 的价值,也避免期望和实际结果之间的落差。


7 Summary

终于把这个最小化的本地 RAG demo 跑通了,没想到从最初梳理流程、整理技术细节,到最后真正完成实操,居然花了我两个多星期。过程中踩了不少坑,比如:下载 HF 模型时因为网速太快导致 safetensors 文件卡住,显存/内存或底层库不兼容导致的 Segmentation Fault,以及 macOS 上没有 CUDA 导致 bitsandbytes 8bit 量化报错……这些问题叠加起来,硬是让我把 rag_query.py 改了几十次~好在,最终还是一个个解决掉了。

不过,凡事有弊也有利,折腾的过程虽然痛苦,但也让我对很多平时不太在意的底层细节有了更直观的理解:比如模型下载的机制、Python 库和硬件加速的耦合关系、内存与显存的边界等等。这些原本“隐形”的东西,在一次次报错和调试中都被迫拆开来看,反而让我收获了意想不到的理解。

通过这次 demo,我对 RAG 的整体流程已经有了更扎实的把握:从文档切片、向量化存储,到检索和重组,再到最后的生成回答,环环相扣,缺一不可。光是搭建出一个能跑通的最小原型,就足够让我体会到 RAG 在“检索 + 生成”这个框架下的力量与局限。它并不是万能的,却是一个非常高效、实用的范式。

下一步,我就可以考虑用 LangChain 或 LlamaIndex(感觉还是需要专门用一篇文章来进行这两者的比较才能梳理清楚)来把这些流程模块化,真正向“生产级”应用靠拢了,目标很明确——和我的博客结合起来,打造一个能为读者提供即时答疑和深度交互的聊天机器人。到那时,遇到的挑战可能会更多:比如如何优化检索效率、如何让答案更贴近上下文、如何平衡准确性和生成的流畅度……但也正是这些挑战,让整个过程变得更有意思。

回过头看,这个 demo 其实就像是一个“试炼场”:它逼着我在有限的资源下想办法解决问题,也让我看到技术的边界与潜力。可以预见,未来在我的博客生态里,RAG 不会是孤立存在的工具,而会成为一块拼图,与知识图谱、智能推荐、自动摘要等能力结合在一起,形成一个更立体的知识系统。

注1:写到这里,我觉得才算是真正完成了“RAG 入门”的第一步,不知道之后还有多少步,也不知道我能走到哪里,只能一步一步来了。

注2:这篇文章实际上去年9月底就写了,不过由于一直给”声音的觉醒”系列文章让路,一让就让了将近半年。

注3:关于文章的排版,一直是让我很蛋痛的事,因为我现在都是直接把obsidian里的markdown格式的文章直接粘贴到文章里,而我用的markdown插件”WP Editor.md”很久没更新了,并且我还用了首段自动缩进的CSS,加上argon主题的一些渲染,所以当文章格式有问题的时候,我都不知道该从哪里开始找问题,毕竟这方面我不擅长,也没啥兴趣研究。也因此导致我最怕复杂的文章——一旦内容种类多了排版看起来就会乱七八糟,就像这篇文章一样~。

📚 系列文章:从零理解 RAG(2 / 2)
12

📌 Content Structure Prompt:
This content belongs to the "AI Learning Map" part, you can view the complete content path from here: AI Learning Map 。
View Related Categories · 3 Matches
📎 Related Articles
Share this article
The blog content is original, please indicate the source when reposting! The RSS address of the blog is:https://blog.tangwudi.com/feed, welcome to subscribe; if needed, you can join theTelegram Groupto discuss questions together.
No Comments

Send Comment Edit Comment


				
|´・ω・)ノ
ヾ(≧∇≦*)ゝ
(☆ω☆)
(╯‵□′)╯︵┴─┴
 ̄﹃ ̄
(/ω\)
∠( ᐛ 」∠)_
(๑•̀ㅁ•́ฅ)
→_→
୧(๑•̀⌄•́๑)૭
٩(ˊᗜˋ*)و
(ノ°ο°)ノ
(´இ皿இ`)
⌇●﹏●⌇
(ฅ´ω`ฅ)
(╯°A°)╯︵○○○
φ( ̄∇ ̄o)
ヾ(´・ ・`。)ノ"
( ง ᵒ̌皿ᵒ̌)ง⁼³₌₃
(ó﹏ò。)
Σ(っ °Д °;)っ
( ,,´・ω・)ノ"(´っω・`。)
╮(╯▽╰)╭
o(*////▽////*)q
>﹏<
( ๑´•ω•) "(ㆆᴗ
😂
😀
😅
😊
🙂
🙃
😌
😍
😘
😜
😝
😏
😒
🙄
😳
😡
😔
😫
😱
😭
💩
👻
🙌
🖕
👍
👫
👬
👭
🌚
🌝
🙈
💊
😶
🙏
🍦
🍉
😣
Source: github.com/k4yt3x/flowerhd
Kaomoji
Emoji
Little Dinosaur
Flower!
Previous
Next
       

👋 Welcome to “Wudi's Personal Blog”

Here, long-term exploration is mainly carried out around the following directions:

🧱 Personal Digital Infrastructure and Blog System Construction
☁️ Cloudflare and Network Architecture Practice
🧠 AI and Knowledge System Exploration
🛡️ Network Security and Access Optimization
🎵 Music and Sound Cognition
👁️ Cognitive Perspectives and Worldviews