Single-agent LLM setups break quickly when your query space is wide. That was the core problem we ran into building Stock Mate, an n8n multi-agent AI inventory automation system for a wholesale distributor.
The client wanted to type questions like "which products are trending down this week?" or "send reorder emails for anything below threshold" and have the system just handle it. Simple ask. The first version we built tried to do it all in one node, one prompt, one model call. It worked for about 40% of queries and confidently hallucinated the rest.
The failure mode was predictable in hindsight. A single Groq-powered node was being asked to both classify the intent of the query and execute the appropriate action in the same step. When the query involved a SQL operation against Supabase, the model would sometimes generate plausible-looking but syntactically wrong SQL. When it involved a Google Trends lookup via SerpAPI, it would make up trend data if the API response was ambiguous.
We weren't dealing with a model quality problem. We were dealing with a scope problem.
The fix was to pull the intent classification step out into its own node, then route to one of three downstream agents based on the result.
The Intent Classifier node receives the raw user query and returns exactly one of three strings: SQL_QUERY, INNER_TREND, or EXTERNAL_TREND. Nothing else. The prompt is constrained deliberately:
You are a router. Classify the user's inventory query into exactly one of:
- SQL_QUERY (direct database lookups, stock counts, product details)
- INNER_TREND (pattern analysis over historical internal sales data)
- EXTERNAL_TREND (market trend lookup requiring external data)
Return only the category label, no explanation.That classification result goes into an n8n Switch node that routes to the appropriate downstream agent. Here is the rough Supabase query structure the SQL agent builds from natural language input (simplified for clarity):
// n8n Function node - SQL agent builds this from the classified query
const userQuery = $input.first().json.query;
const systemPrompt = `You generate PostgreSQL SELECT statements for a product inventory schema.
Tables: products(id, name, sku, stock_qty, reorder_threshold), sales(product_id, qty, date).
Return only valid SQL, no explanation.`;
const sql = await callGroq(systemPrompt, userQuery);
const result = await supabase.rpc('run_query', { query: sql });
return result;The External Trend agent uses SerpAPI to fetch Google Trends data for the product category, then passes that raw data to a separate summarization prompt. Keeping the fetch and the summarization in two separate nodes meant we could catch bad API responses before they contaminated the summary.
The third output, beyond queries, was automated reorder emails. When the SQL agent's response showed stock_qty < reorder_threshold, a downstream condition node triggered a Gmail API call to the assigned vendor.
This part actually worked cleanly from the start. The logic was simple enough that a single node could handle it without a complex prompt. Condition nodes, not LLMs, handle branching logic that can be expressed as a comparison. That's worth keeping in mind when you're designing these flows.
The intent classifier originally had five categories. We had split "trend analysis" into "weekly trend," "monthly trend," and "category trend." In practice this made the classifier more confused, not less, and we were collapsing them back together in every downstream agent anyway. Three categories was the right number for this dataset and query type. Don't over-engineer the routing layer.
And we spent a week trying to make one universal agent handle all three types with a very long system prompt before admitting it wasn't working. The switch to specialized agents happened in about two days and fixed the hallucination rate almost entirely.
Stock Mate now runs as an always-on n8n workflow. The client types natural language queries into a basic web form. The classifier routes, the appropriate agent executes, and results come back as formatted summaries. Routine reorders go out automatically without any human review.
If you are building something similar and hitting the same single-agent hallucination problems, splitting intent classification into its own constrained node is the fastest fix to try. It's a small structural change and the improvement is usually immediate.