How I made AI Search “Whatever data you want to see…” for Find Me in Chicago.
Building Text-to-SoQL was a journey in pragmatic AI especially around the CHAD challenges:
- managing token Cost through minimal architectures
- suppressing Hallucination with strict schema and dialect grounding
- converting human Ambiguity into precise intent
- taming hardware non-Determinism through regex normalization and snapshot testing
AI is not yet a replacement for software engineering, and having a probabilistic engine in your pipeline is like playing with fire. Dangerous but also very very fun.
Traditional Text-to-Sql LLM narratives talk about heavy agentic frameworks, vector databases, and multi-turn verification loops for structured translation. In reality, for my Text-to-SoQL application, a single-shot prompt wrapped in a deterministic sandwich of serverless C#, regex macros, and schema grounding gave me great results at a minimal cost.
One thing I learned is you can’t fight LLM non-determinism, just learn to engineer around it. Accepting that the same input will only randomly return the same output is not entirely comfortable, but having worked so long in production systems, with long remediating “heisen-bugs”, it works on my machine, and we don’t have the logs, it is not a big stretch. Working with LLMs requires the same mindset shift any developer goes through in their journey to support complex and critical requirements. You don’t need mathematical perfection; you need bounded, reliable utility.
All of this engineering exists to serve a real civic mission. The Chicago Data Portal is a treasure trove of data, but raw REST APIs, and uninviting user interfaces keep it away from the people who care. The true measure of my Text-to-SoQL journey won’t be if it achieves a good score on an academic benchmark. The true measure my Text-to-SoQL implementation is if it lets someone, anyone, in Chicago, or wherever, type a messy, typo-ridden question into their phone and immediately get answers about the community.