Tools or RAG: when to use each
On one project I used tools to query data that changed all day; on my site I use RAG over text that barely changes. What decides between them, and when it's worth combining both.
The two AI systems I know best solve the same problem in opposite ways. In the CEDAE case, the AI answered questions about more than 10,000 inspections by calling ready-made functions, tools. In Nox, this site's chat, it answers questions about me by reading passages picked by a search step, RAG. Either way the model needs information that wasn't in its training data. The difference is what kind of information that is.
Both mechanisms, one sentence each
- Tools (function calling): you describe functions to the model, it decides which one to call and with which arguments, your code runs it and returns the result. The model asks; you fetch.
- RAG (retrieval-augmented generation): before calling the model, you search your content for the passages closest to the question and put them in the prompt. The model only reads what you picked.
When tools win
At CEDAE the most common questions kept repeating: who owns this property, which technician did this inspection, is there a duplicated address, what's the payment status. The data behind them changed several times a day and the answer had to be exact. RAG would have been the wrong tool: indexing a table that changes all the time means reindexing nonstop, and similarity search returns what's close, not what's right. "Roughly the owner" doesn't cut it.
With tools, each repetitive question became a ready-made function. The model only picked the function and filled in the arguments; some of them carried their own SQL query. Free-form queries were left for whatever fell outside the pattern. I got three things out of it: exact answers, because they come from the database at that moment; less room for the model to make things up, because it doesn't build the query from scratch; and control, because each function does only what I wrote. All of it ran against a separate, read-only database, so no tool could change data, by permission rather than by asking nicely in the prompt.
When RAG wins
Nox answers questions about a small body of prose that changes when I publish something: background, case studies, stack, notes. The questions are open-ended ("why did he leave law?") and almost never map to an exact field. That's where RAG fits: I embed the site at build time, run a similarity search on each question and hand over the closest passages. No table, no query, and the model has no power at all: it only reads.
Its weak spot showed up in the evaluation: the content was written in first person and the questions came in third, and the search didn't bring them together. The fix was giving every passage a third-person title. That kind of problem only exists in RAG, because it depends on how the text is split and labeled.
The rule of thumb I use now
- Does the data change often? Tools, querying the source at answer time.
- Does the answer need to be exact (a number, a name, a status)? Tools.
- Do questions repeat in a predictable shape? A ready-made tool for each.
- Is the content prose that rarely changes, with open-ended questions? RAG.
- Will the model be able to act? Then the real question is the least power it can get away with, and that applies to both.
And when to combine them
The interesting cases need both. At CEDAE, today I'd put RAG over what is text (inspection summaries, operating manuals) and keep tools for structured data. I'd also turn retrieval itself into a tool, so the model decides when to search instead of me always searching. The next step is exposing those tools through a standard protocol, MCP, so any assistant can use them. I did that on this site: the same information Nox uses is available as tools at /api/mcp.