An embedding is a list of numbers (a vector) that captures what a piece of text means. Texts with similar meanings get similar vectors, even when they use different words. “How do I cancel my plan?” and “stop my subscription” end up close together.
This is how a knowledge base finds the right passages for a question: it turns the question into a vector and returns the passages whose vectors are closest.
Where embeddings are used in AWFlow
Section titled “Where embeddings are used in AWFlow”flowchart LR Source["Page, file or text"] --> Indexer["Indexer"] Splitter["Text splitter"] --> Indexer Indexer --> KB["Local Knowledge (knowledge base)"] Question["Question"] --> Agent["RAG / Q&A / Tools Agent"] KB --> Agent Agent --> Answer["Answer with sources"] class KB awf-data
- The Indexer splits text into chunks (with a text splitter) and saves each chunk with its vector in a knowledge base.
- Local Knowledge connects that knowledge base to an agent.
- The RAG Agent, Q&A Agent or Tools Agent searches it before answering.
You can also fill a knowledge base without a workflow: on its page, click Add sources to add files, web pages, open tabs or pasted text, or save from anywhere.
Embedding models
Section titled “Embedding models”The embedding model turns text into vectors. Every knowledge base remembers the model it was built with and always uses it, because vectors from different models can’t be compared.
| Model | Where it runs | Node |
|---|---|---|
| all-MiniLM-L6-v2 (default) | In your browser, free, private | Local Embeddings |
| Ollama models | On your computer | Ollama Embeddings |
| OpenAI | Cloud, needs your API key | OpenAI Embeddings |
| Google (Gemini) | Cloud, needs your API key | Google Embeddings |
| Cohere | Cloud, needs your API key | Cohere Embeddings |
| Voyage AI | Cloud, needs your API key | Voyage Embeddings |
| Hugging Face | Cloud, needs your API key | Hugging Face Embeddings |
A new knowledge base uses the on-device model unless you connect another embeddings node on its first run. To move a knowledge base to another model later, use Re-embed with… (see Change the search engine).
Meaning and keywords
Section titled “Meaning and keywords”Pure vector search is good at meaning but can miss exact words, names and codes, such as an invoice number. The agents’ Match by option controls this:
- Keywords and meaning (default) also finds exact words.
- Meaning only uses vector similarity alone.
Other search options (Search type, Minimum relevance, Metadata filter) are described in Knowledge base options.
Getting good results
Section titled “Getting good results”- Chunk size matters. Small chunks find precise facts; large chunks keep more context. Start with the splitter’s defaults and adjust if answers miss details or lack context.
- Index clean text. Navigation menus, cookie banners and footers add noise. Extract the main content first, for example with Get All Text on the right part of the page.
- Name your sources. Set Source name and Source link on the Indexer so answers can cite them, and so re-indexing replaces the old version instead of adding a copy.
- Retrieve a few passages, not many. The agents use 4 by default. More passages give the model more to read, but also more to get confused by.