Daniel Valle
LLM reliability engineer

I fix LLM products that work in the demo and break in production.

Reliability, evals, cost and latency for AI agents. Since 2025 I have built and run Flora, a multi-agent voice companion for seniors, in production: about 15 repositories, 18 containers, DeepSeek and Mistral models with host failover, EU-hosted. Most of what I fixed returned HTTP 200.

~6.5 s → 1.26 sclassifier p50 in production, after a model change
−61% tokenstypical prompts, from where one parameter was placed
0% → 99%prompt cache hits, after fixing our client layer

Services

Scoped, async, measured before and after. 10 to 15 hours a week, evenings CET, which overlaps US business hours.

Production agent build

An agent for your use case, built the way I built Flora: labelled test cases from day one, a safe fallback on every model call, tracing, and cost and latency measured on your prompts. WhatsApp and voice channels, Portuguese (PT and BR) and English. Phased, part-time delivery.

fixed price per phase, from $2,500

LLM production audit

A written report on your pipeline: silent failures, cost leaks, latency tail, eval gaps, and a prioritised fix list. Every finding measured.

fixed price, from $1,500

Cost and latency reduction

Before and after numbers on your real prompts. Token budgets, caching that actually hits, provider and model choice, request shape.

fixed scope or hourly

Evals, debugging and migration

An eval harness with a measured noise floor, fixes for unexpected agent behaviour and structured-output failures, or a move to EU-hosted models with measured parity and safe fallbacks.

fixed scope or hourly

Case studies

All measured on Flora. Each result depended on Flora's situation and the models available to it at the time, so read them as a method, not a promise of the same number on another system. Sample sizes are given; estimates are labelled as estimates.

LATENCY

A classifier on every message: production p50 from about 6.5 s to 1.26 s

Context

Every user message goes through an intent and mood classifier before Flora starts replying. It ran on Mistral Small, at a production median of about 6.5 s.

Diagnosis

The call was decode-bound: its time tracked output length. The prompt also asked for a chain-of-thought inside the JSON, and on the candidate model that reasoning was written after the decision fields, so it could not inform them.

What I did

Moved this one classifier to DeepSeek V4 Flash on an EU host, removed the chain-of-thought for that model only (Mistral writes fields in schema order, so it keeps it there), pinned a strict output schema, and kept Mistral as a fallback with an alert. Validated on a 130-case labelled set at n=5, re-running borderline cases.

  • Labelled set: p50 2,409 ms to 589 ms. Most of that came from the faster model; removing the unused chain-of-thought cut a further ~36%. Output went from about 825 to about 230 tokens.
  • Accuracy: 7 cases fixed, 1 regressed (a pacing case, accepted after analysis); all safety cases stable.
  • Honest limit: this worked because a faster model of equal quality existed for this task. Prompt caching was not the lever here; reading the prompt was only ~370 ms of the call.
~6.5 s1.26 sproduction p50 · 0 fallbacks over the first ~60 turns
COST · LATENCY

Where one parameter was placed: −61% output tokens

Context

Flora's reply model (DeepSeek V4 Flash) was getting slower and hitting its timeout. Nothing errored.

Diagnosis

The reasoning-effort parameter was sent at the top level of the request. On our provider, that placement produced much longer generations than the same value nested where the model's chat template reads it. Only the token counts showed it.

What I did

An A/B on the request shape with real production prompts, interleaved, then three quality checks.

  • One heavy prompt showed a bigger gap (12,429 to 3,259 tokens, n=5 per arm). It was an outlier, so the typical figure below is the one that counts.
  • Quality: persona suite 15/15 on both arms, reasoning suites 25/29 vs 24/29, blind comparison on 11 real conversations 7 to 4. That is no detectable regression, not an improvement.
  • Part of the gain is that the model writes its long self-check less often. I track that rate as the risk signal.
5,865 tok · 21.5 s2,296 tok · 9.1 stypical prompts, n=12 per arm: −61% tokens, −57% latency
SILENT FAILURES · COST

Evaluating a gateway: settings accepted, not applied

Context

We evaluated routing the classifier stack through an LLM gateway (one key, many providers).

What I measured

  • Prompt caching. Direct, 2,288 of 2,311 prompt tokens came from cache on repeat calls. Through the gateway, three identical calls reported identical cost, so the cache key was accepted but not applied. On a path making 12 to 15 calls per user turn, I estimated roughly 5x the input cost (from cache pricing, not from a bill).
  • Latency. n=20 per arm, interleaved, through the real classifier: p50 4,813 ms direct vs 6,232 ms via the gateway (+29%).
  • Reasoning. On another gateway route, 12 different request shapes all returned no reasoning, although the catalog listed reasoning support. A catalog flag is not evidence; a live trace is.

Decision, and a fix in our own code

Kept the direct route. Separately, a client-library layer in our code was overwriting the cache key before sending; after fixing it, cache hits on that call went from 0% to 99%. The probing method became llm-wire-check.

0% cached99% cached4,030 of 4,072 prompt tokens, after the client-layer fix
MIGRATION

Moving off OpenAI to EU-hosted models

Context

Flora talks with seniors about personal matters. The goal was no US model vendor in that path, without losing quality on safety-critical flows: medication reminders, consent, destructive task edits.

What I did

Migrated every GPT call in phases with one rule: extraction and classification to Mistral Small, text generation to DeepSeek V4 Flash. Every call site got a safe fallback when the model fails (block, skip, or keep the original). The reply model also fails over across a chain of hosts.

  • Supervisor router on Mistral Small: p50 379 ms, p95 465 ms.
  • OpenAI fully removed: no GPT model, key or identifier left.
  • Honest limit: the host fallback covers a host failing. When the shared upstream model slows down, some turns still time out; that is still open.
OpenAIEU-hostedadversarial routing 29/30 on the new model (12/12 on the original set); the previous model had routed abusive messages into the task pipeline
EVALS

Making the eval suite tell the truth

Context

The test suite reported regressions that were not there, and passes that were not real.

What I found and fixed

  • 32 suites, every safety suite included, failed all their cases because of one stale mock. Wiring, not a regression.
  • A suite passed 11/11 with zero real model calls: a missing API key made the classifier return its default, so every "expect false" assertion passed vacuously.
  • Measured the harness against itself: two runs of an identical prompt disagreed on 5 cases at n=3, and at n=9 borderline cases still spread 0.44 against a 0.34 regression threshold. Added confirmation re-runs.
9/9blockedthe harness caught a faster prompt that took two safety cases to 0/9. It never shipped.
LATENCY · OBSERVABILITY

The tracing tool was a large part of the overhead

Context

Every turn paid non-model framework time: p50 3.83 s, p90 10.56 s over 37 turns.

Diagnosis and fix

The largest single cause was the tracing callback: it re-serialised the whole conversation state about 44 times per turn. I summarised nested payloads and deferred the work, keeping full traces readable. Measured in a benchmark; no production p50 was recorded after the fix.

+679 to 930 ms+13 to 32 mshandler cost at 400 messages of history · sub-graph 777 ms → 402 ms (407 ms with tracing off)

Open source

llm-wire-check sends controlled request variants to any OpenAI-compatible endpoint and reports what actually changed: reasoning switches that are ignored, effort levels that inflate output, caches that never hit, models silently substituted. Because almost every one of these failures returns 200.

About

I am Daniel Valle, a software engineer based in Portugal. I build Flora, a voice companion for seniors who live alone: she talks with them every day, in three Portuguese dialects, over WhatsApp and an Android app, and turns those conversations into follow-up for the people who care for them.

The backend is a TypeScript and NestJS system with a LangGraph supervisor routing to task, preference and companion agents, voice notes through speech-to-text and text-to-speech, and an LLM layer spread across DeepSeek and Mistral with automatic failover. I care about measured claims: every number on this page has a sample size behind it.

Languages: Portuguese (native, PT and BR), English (fluent).

Contact

Tell me what is slow, expensive or unpredictable. I will reply within a day.