Skip to content
FlowHubFluxonLab
RAG & Knowledge Basesfree

RAG Chatbot with Small Local Llms (Ollama) — No Tool Calling

by Wassim AbidUpdated Aug 2026
RequiresOllama
Share Post Share
SOSplit Sub-QueriesSplit Sub-Queri…AgAggregate Matching ChunksAggregate Match…AgAggregate All Retrieval ResultsAggregate All R…IfAny chunk?SeClean RAG outputFiKeep score over 0.4Keep score over…SeSay no chunk matchSay no chunk ma…SePrepare loop outputPrepare loop ou…MPPostgres Chat Memory (Small Talk)Postgres Chat M…CoRemove Think Tags (RAG Path)Remove Think Ta…SwSwitchCoJSON FormatterAgSmall Talk AI AgentSmall Talk AI A…Ollama Chat Model (Small Talk — Qwen3:14b)Ollama Chat Mod…Ollama Chat Model (Classifier — Qwen2.5:7b)Ollama Chat Mod…MPPostgres Chat Memory (RAG Answer)Postgres Chat M…AgAnswer Generator AI AgentAnswer Generato…Ollama Chat Model (Answer Generator — Qwen3:14b)Ollama Chat Mod…SILoop Over Sub-QueriesLoop Over Sub-Q…EOOllama Embeddings (BGE-M3)Ollama Embeddin…VSPGVector Store — Retrieve ChunksPGVector Store …CoRemove Think Tags (Small Talk Path)Remove Think Ta…WeWebhookRTRespond to WebhookRespond to Webh…CLUnderstand RequestUnderstand Requ…123456789101112131415161718192021222324252627
1/5
STEPS · 27
Starts on an incoming request

On a webhook message, retrieves context from a PGVector store and answers with local Ollama LLMs and embeddings, replying via the webhook.

Tags

webhookadvancedOllamaInternal WikiAI RAGdiscoveredpending-review
Connects
Ollama
CategoryRAG & Knowledge Bases
Triggerwebhook
Complexityadvanced
Nodes25
AddedJun 27, 2026