How I made AI Search “Whatever data you want to see…” for Find Me in Chicago.
In any engineering project cost is an important constraint. Arguably, it is the most important constraint. One thing we all know about AI inference is that it is expensive. AI requires a lot of compute, high tech manufacturing, water, electricity, networking, and thus money to be effective. To minimize costs, the first things to look at are token consumption, cloud infrastructure, and engineering complexity.
AI tokens are not all priced equally. Across AI providers, the difference between the lowest cost model and the highest cost model can easily be an order of magnitude per token. This is not necessarily an apples to apples comparison. The more expensive tokens are theoretically “better” for certain workloads and might be a better value. So examine your workload to decide which tokens are best.
For Find Me in Chicago, I had already decided on Google Gemini as my AI LLM provider and ran basic tests to determine the most cost effective model for quality results. I was surprised to find (in my not very thorough testing), that the lowest cost model seemed to provide the highest quality results. I assume the reason is a mix of decreased Hallucination and increased Determinism or maybe I just didn’t measure it right.
Cloud infrastructure is often cheap compared to AI tokens. Any work I can move off my AI prompt and into my cloud infrastructure is going to save money.
The trade-off is moving things from the AI prompt to the cloud infrastructure increases engineering complexity. Engineers are not cheap either. Engineering complexity also compounds over time with increased maintenance, debugging, and operational overhead once development is complete. If you save a penny per request, how many requests will it take before you make back all the time and money you spent optimizing the request. As Donald Knuth said: “Premature optimization is the root of all evil.”
For Text-to-SoQL, the cost isn’t just about reducing the number of tokens, or choosing the cheapest token. It is about finding the right balance between the AI LLM, the cloud infrastructure, and the engineering effort. So I started with the simplest thing that could possibly work. The cheapest tokens, the simplest infrastructure and the least complex engineering. In practice, this meant using the latest Google Gemini Flash Lite model available at the time, serverless functions, and following the YAGNI principle of “You Ain’t Gonna Need It.”
In practice that meant I start with no RAG (via a vector database or whatever), just one shot with no agentic feedback loops or query verifications. It either works to “Just tell me the data you want to see…” or it doesn’t.