Text-to-SoQL via AI – Determinism – The Butterfly Effect

How I made AI Search “Whatever the data you want to see…” for Find Me in Chicago.

I started my Text-to-SoQL journey with a naive misconception that cranking down these settings would make my queries mostly deterministic:

  • temperature: 0 – sharpen the LLM to always choose the highest probability token
  • topK: 1 – forces the model to evaluate only the single top-ranked token even though that itself is non-deterministic
  • seed: 1 – lock the randomness, but with temperature: 0 and topK: 1, it basically does nothing

In general, a low temperature will improve determinism. But even at zero, it does not guarantee identical output. There is nothing you can do to ensure full determinism from an LLM so the same input always results in the same output. This is not a bug. LLM inference introduces non-determinism through sampling, numerical computation, and server infrastructure.

LLMs face a cascading series of non-deterministic computations and they start in the hardware. The result is commonly known as The Butterfly Effect, where even the smallest action can sometimes initiate massive consequences.

  • Floating Point Noise: Floating-point addition on GPUs is imprecise and non-associative.
  • Distributed Inference: Queries hit different GPUs with varying memory layouts, competing workloads, and thermal states. Yes Virginia, this means the weather can affect your AI model.
  • Dynamic Batching: Server batch sizes shift unpredictably, altering thread block shapes and reduction tree orders.
  • KV-Cache Quantization: Compressing memory from 32-bit floats to 4-bit/8-bit ints introduces cumulative rounding errors.
  • Inference Optimizations: LLMs may reorder operations for speed, and speculative decoding introduces branching choices.

LLMs autoregressively generate text one token at a time. Each token’s probability distribution is conditioned on *every* token that came before it. If floating-point noise causes the model to pick a slightly different token at position 15 (say, emitting an extra space or choosing `LIKE` before `AND`), then:

  • Token 16 now sees a different history
  • Its probability distribution shifts and it might pick a different token
  • Token 17 sees an even more different history
  • Rinse and repeat…

I started with the smallest GCP model I thought could possibly work, Gemini Flash-Lite. It was originally chosen because of cost, but, according to Gemini, smaller models can help with determinism with caveats making it a complex trade off.

The case for smaller models:

  • Less Creativity: Less likely to overthink prompts. Lower parameter counts mean less capacity for unprompted contextual embellishments or synonym swapping.
  • Fewer Competing Paths: A narrower distribution of query patterns means fewer choices at equal probabilities—fewer tie-breaks and flipped tokens when subtle float variations occur.
  • Simpler Hardware Layout: They fit on fewer GPUs, cutting floating-point noise at the source.

The case for larger models:

  • Less Aggressive Quantization: Larger models often have higher numerical precision. This makes them less dependent on FP8/INT8 quantization to hit low latency targets, keeping the baseline noise floor low.
  • Attention Redundancy: Larger models have more heads and layers to average out micro-variations.
  • Decisive Margins: Larger models develop decisive logit preferences, pulling rank #1 far enough ahead to resist noise.

Gemini basically wrote these points critiquing determinism for its own models. If any Gemini engineers out there have constructive feedback on its accuracy, let me know.

My understanding of it all is that you trade cluster-level hardware drift in exchange for razor-thin logit margins. The hardware noise is smaller, but so is the margin of error. For now I am sticking with the path of lowest cost, which is the smaller model.

So how do you test this non-determinism? Even with whitelisting, constraint enforcement, and a well defined output schema, fields show up in different orders, semantically equivalent number and date operations show up, spurious parentheses are often added at random, and sometimes an extra order by field will show up that doesn’t really affect the results.

I opted for snapshot tests with multiple verified outputs as a human-in-the-loop substitute for the non=determinism. I don’t have a full SoQL expression parser to understand syntax operator precedence but I did write a naive normalization routine with RegEx to reduce the number of variations I need to accept.

That is my not ideal but pragmatic solution for now to control the chaos.

Send your comments to me on LinkedIn.