
How I made AI Search “Whatever data you want to see…” for Find Me in Chicago.
LLMs are trained on billions of lines of SQL. SoQL looks like SQL but differs in critical ways:
- No `CASE`/`WHEN` keywords, although SoQL does support a case() function
- No subqueries
- No `JOIN` supported (for the public endpoint used)
- Different function definitions, especially for dates (e.g. `date_extract_y` vs `YEAR()`)
LLMs appear to have a strong instinct to emit valid SQL. Every SoQL-specific constraint in a prompt is fighting against the model’s massive SQL training. So you have to be very clear about what is and what is not valid SoQL. I set this in the system prompt before any Schema Grounding. Some day the frontier models will hopefully get SoQL support if you just ask for it, but for now SoQL grounding is manual and consumes tokens on every request.
The grounding also defines intent and interpretation binding common aggregate keywords (“how many”, “summarize”) to count(*) and ensures grouping and having are correctly formatted. The SoQL global prompt also specifies a strict JSON serialization for the output format.
Like SQL, SoQL has many equivalent ways to generate the same logical expression. The more restrictions you place on the LLM, the more deterministic is the output. Hence, I also restrict use of `AS` for aliasing.