My data engineering knowledge base started as straight RAG. Every question followed the same path: embed the query, pull the nearest chunks from the vector store, hand them to the model, return an answer. It worked for lookups. It fell apart on everything else.
A fixed retrieval path retrieves whether it should or not. Ask it to write a SQL query and it would still go digging through documentation chunks first. Ask it something current and it had nothing, because the only knowledge it could reach was whatever I had indexed. The pipeline was making the decision, not the model.
So I rebuilt it as an agent using OpenAI function calling. Now the model decides whether a tool is needed at all, which one, and in what order. It has three: knowledge base retrieval over Pinecone and LlamaIndex, SQL generation that is aware of the target dialect, and web search through Tavily for anything outside the index. A simple question gets answered directly. A harder one gets the right tool, and the model can chain several across multiple reasoning rounds before it responds.
It also keeps conversation memory across turns, so follow-up questions build on what came before instead of starting cold.
The part I care about most is failure handling. When a tool errors, I return the error back to the model rather than dropping the request. It reads what went wrong and tries another route. A failed SQL call or an empty search does not end the conversation, it just becomes the next thing the model reasons about. That one decision made the whole system feel reliable instead of brittle.
Same models, same vector store, same data. The difference is who is deciding. Handing that choice to the model instead of hard-coding it in a pipeline is the line between a retrieval script and something that can actually work a problem.
Live demo → View code →