LLM reliability engineer
I fix LLM products that work in the demo and break in production.
Reliability, evals, cost and latency for AI agents. Since 2025 I have built and run Flora, a multi-agent voice companion for seniors, in production: about 15 repositories, 18 containers, DeepSeek and Mistral models with host failover, EU-hosted. Most of what I fixed returned HTTP 200.
~6.5 s → 1.26 sclassifier p50 in production, after a model change
−61% tokenstypical prompts, from where one parameter was placed
0% → 99%prompt cache hits, after fixing our client layer
Case studies
All measured on Flora. Each result depended on Flora's situation and the models available to it at the time, so read them as a method, not a promise of the same number on another system. Sample sizes are given; estimates are labelled as estimates.
LATENCY
A classifier on every message: production p50 from about 6.5 s to 1.26 s
Context
Every user message goes through an intent and mood classifier before Flora starts replying. It ran on Mistral Small, at a production median of about 6.5 s.
Diagnosis
The call was decode-bound: its time tracked output length. The prompt also asked for a chain-of-thought inside the JSON, and on the candidate model that reasoning was written after the decision fields, so it could not inform them.
What I did
Moved this one classifier to DeepSeek V4 Flash on an EU host, removed the chain-of-thought for that model only (Mistral writes fields in schema order, so it keeps it there), pinned a strict output schema, and kept Mistral as a fallback with an alert. Validated on a 130-case labelled set at n=5, re-running borderline cases.
- Labelled set: p50 2,409 ms to 589 ms. Most of that came from the faster model; removing the unused chain-of-thought cut a further ~36%. Output went from about 825 to about 230 tokens.
- Accuracy: 7 cases fixed, 1 regressed (a pacing case, accepted after analysis); all safety cases stable.
- Honest limit: this worked because a faster model of equal quality existed for this task. Prompt caching was not the lever here; reading the prompt was only ~370 ms of the call.
~6.5 s1.26 sproduction p50 · 0 fallbacks over the first ~60 turns
COST · LATENCY
Where one parameter was placed: −61% output tokens
Context
Flora's reply model (DeepSeek V4 Flash) was getting slower and hitting its timeout. Nothing errored.
Diagnosis
The reasoning-effort parameter was sent at the top level of the request. On our provider, that placement produced much longer generations than the same value nested where the model's chat template reads it. Only the token counts showed it.
What I did
An A/B on the request shape with real production prompts, interleaved, then three quality checks.
- One heavy prompt showed a bigger gap (12,429 to 3,259 tokens, n=5 per arm). It was an outlier, so the typical figure below is the one that counts.
- Quality: persona suite 15/15 on both arms, reasoning suites 25/29 vs 24/29, blind comparison on 11 real conversations 7 to 4. That is no detectable regression, not an improvement.
- Part of the gain is that the model writes its long self-check less often. I track that rate as the risk signal.
5,865 tok · 21.5 s2,296 tok · 9.1 stypical prompts, n=12 per arm: −61% tokens, −57% latency
SILENT FAILURES · COST
Evaluating a gateway: settings accepted, not applied
Context
We evaluated routing the classifier stack through an LLM gateway (one key, many providers).
What I measured
- Prompt caching. Direct, 2,288 of 2,311 prompt tokens came from cache on repeat calls. Through the gateway, three identical calls reported identical cost, so the cache key was accepted but not applied. On a path making 12 to 15 calls per user turn, I estimated roughly 5x the input cost (from cache pricing, not from a bill).
- Latency. n=20 per arm, interleaved, through the real classifier: p50 4,813 ms direct vs 6,232 ms via the gateway (+29%).
- Reasoning. On another gateway route, 12 different request shapes all returned no reasoning, although the catalog listed reasoning support. A catalog flag is not evidence; a live trace is.
Decision, and a fix in our own code
Kept the direct route. Separately, a client-library layer in our code was overwriting the cache key before sending; after fixing it, cache hits on that call went from 0% to 99%. The probing method became llm-wire-check.
0% cached99% cached4,030 of 4,072 prompt tokens, after the client-layer fix
MIGRATION
Moving off OpenAI to EU-hosted models
Context
Flora talks with seniors about personal matters. The goal was no US model vendor in that path, without losing quality on safety-critical flows: medication reminders, consent, destructive task edits.
What I did
Migrated every GPT call in phases with one rule: extraction and classification to Mistral Small, text generation to DeepSeek V4 Flash. Every call site got a safe fallback when the model fails (block, skip, or keep the original). The reply model also fails over across a chain of hosts.
- Supervisor router on Mistral Small: p50 379 ms, p95 465 ms.
- OpenAI fully removed: no GPT model, key or identifier left.
- Honest limit: the host fallback covers a host failing. When the shared upstream model slows down, some turns still time out; that is still open.
OpenAIEU-hostedadversarial routing 29/30 on the new model (12/12 on the original set); the previous model had routed abusive messages into the task pipeline
EVALS
Making the eval suite tell the truth
Context
The test suite reported regressions that were not there, and passes that were not real.
What I found and fixed
- 32 suites, every safety suite included, failed all their cases because of one stale mock. Wiring, not a regression.
- A suite passed 11/11 with zero real model calls: a missing API key made the classifier return its default, so every "expect false" assertion passed vacuously.
- Measured the harness against itself: two runs of an identical prompt disagreed on 5 cases at n=3, and at n=9 borderline cases still spread 0.44 against a 0.34 regression threshold. Added confirmation re-runs.
9/9blockedthe harness caught a faster prompt that took two safety cases to 0/9. It never shipped.
LATENCY · OBSERVABILITY
The tracing tool was a large part of the overhead
Context
Every turn paid non-model framework time: p50 3.83 s, p90 10.56 s over 37 turns.
Diagnosis and fix
The largest single cause was the tracing callback: it re-serialised the whole conversation state about 44 times per turn. I summarised nested payloads and deferred the work, keeping full traces readable. Measured in a benchmark; no production p50 was recorded after the fix.
+679 to 930 ms+13 to 32 mshandler cost at 400 messages of history · sub-graph 777 ms → 402 ms (407 ms with tracing off)
About
I am Daniel Valle, a software engineer based in Portugal. I build Flora, a voice companion for seniors who live alone: she talks with them every day, in three Portuguese dialects, over WhatsApp and an Android app, and turns those conversations into follow-up for the people who care for them.
The backend is a TypeScript and NestJS system with a LangGraph supervisor routing to task, preference and companion agents, voice notes through speech-to-text and text-to-speech, and an LLM layer spread across DeepSeek and Mistral with automatic failover. I care about measured claims: every number on this page has a sample size behind it.
Languages: Portuguese (native, PT and BR), English (fluent).