Hamburg container terminal cranes receding into morning haze

Scaling Applied AI Applications to Production

Some learnings about building Non-Deterministic Systems in Production and what Engineering Teams should not do.

Most teams know how to ship deterministic software. You write a function, you write a test, the test either passes or it doesn't, and you sleep fine. The contract between you and your code is clear: same input, same output, every time.

Applied AI breaks that contract.

The moment you put an LLM, an embedding model, or any probabilistic component into a production path, you've taken on a system that can return different answers to the same question on Tuesday versus Wednesday. That isn't a bug to be patched out. It's the substrate. And yet most engineering practices — CI, code review, monitoring, on-call playbooks — were built for a world where that doesn't happen.

This is what I keep seeing engineering teams underestimate when they move from a working demo to something a paying customer depends on. Below are the patterns we've found useful at HUBBLR, both in our own products and in the AI work we do for clients.

Stop treating the model as the system

The first mistake is conflating "the model" with "the product." The model is one component. The product is everything around it: the prompt construction, the retrieval layer, the validation, the fallback paths, the human-in-the-loop, the observability, the cost controls, and the place where the output actually lands in a workflow.

When something goes wrong in production, it is almost never the model that's wrong — it's the scaffolding around the model. A worse model with a thoughtful pipeline beats a frontier model wired up naively, every time.

So when you scope an applied AI feature, scope the scaffolding first. The model is the easy part to swap out later.

Evaluations are your new test suite

Unit tests assume determinism. Evals don't.

An eval is a dataset of representative inputs paired with either expected outputs, scoring criteria, or both. You run your pipeline against it whenever something changes — a new prompt, a new model version, a new retrieval step — and you watch the aggregate scores. Not pass/fail. Distribution.

This sounds obvious until you try to put it into a CI pipeline. Then you have to answer real questions: what's the regression threshold? Who owns the eval set? How do you keep it from going stale? How do you keep it from being a proxy for the demo your sales team likes? An eval set that only contains the happy path is worse than no eval set, because it gives you false confidence.

The teams that get this right treat their eval suite as a living artifact, owned jointly by engineering and whoever holds the domain knowledge. New failure modes from production get added back. Old ones don't get deleted just because they're embarrassing.

A minimal pattern we use to wire evals into CI looks like this:

# tests/evals/test_classifier_eval.py
import json
import pytest
from pathlib import Path
from app.pipeline import classify_product

EVAL_SET = json.loads(Path("evals/products.json").read_text())
REGRESSION_THRESHOLD = 0.92  # don't merge below this

@pytest.mark.eval
@pytest.mark.parametrize("case", EVAL_SET, ids=lambda c: c["id"])
def test_classification_case(case, eval_recorder):
    result = classify_product(case["input"])
    correct = result.category == case["expected_category"]
    eval_recorder.record(
        case_id=case["id"],
        passed=correct,
        predicted=result.category,
        expected=case["expected_category"],
        confidence=result.confidence,
    )

def test_aggregate_accuracy(eval_recorder):
    # Runs after the parametrized cases above.
    accuracy = eval_recorder.pass_rate()
    assert accuracy >= REGRESSION_THRESHOLD, (
        f"Eval accuracy {accuracy:.2%} below threshold "
        f"{REGRESSION_THRESHOLD:.2%}. Inspect regressions before merging."
    )
# tests/evals/test_classifier_eval.py
import json
import pytest
from pathlib import Path
from app.pipeline import classify_product

EVAL_SET = json.loads(Path("evals/products.json").read_text())
REGRESSION_THRESHOLD = 0.92  # don't merge below this

@pytest.mark.eval
@pytest.mark.parametrize("case", EVAL_SET, ids=lambda c: c["id"])
def test_classification_case(case, eval_recorder):
    result = classify_product(case["input"])
    correct = result.category == case["expected_category"]
    eval_recorder.record(
        case_id=case["id"],
        passed=correct,
        predicted=result.category,
        expected=case["expected_category"],
        confidence=result.confidence,
    )

def test_aggregate_accuracy(eval_recorder):
    # Runs after the parametrized cases above.
    accuracy = eval_recorder.pass_rate()
    assert accuracy >= REGRESSION_THRESHOLD, (
        f"Eval accuracy {accuracy:.2%} below threshold "
        f"{REGRESSION_THRESHOLD:.2%}. Inspect regressions before merging."
    )

The point isn't the framework — it's the shape. Per-case results stored for later inspection, an aggregate threshold that gates merges, and a structure you can extend with latency, cost, or LLM-as-judge scores without rewriting the harness.

Logging isn't observability

In a deterministic system, a log line tells you what happened. In a non-deterministic system, a log line tells you what one specific roll of the dice looked like. That's not nothing — but it's not enough.

You need three things instead:

The full trace of the request, including the exact prompt sent, the exact response received, the retrieved context, the tool calls, and the user-visible output. Not summaries — the raw payloads. Storage is cheap; debugging a hallucination six weeks later without the original prompt is expensive.

Aggregate metrics that tell you what the distribution of behavior looks like over time. Average latency, p95 latency, token usage, cost per request, refusal rate, tool-call success rate. These shift before anything breaks loudly.

Feedback loops from the user. A thumbs up/down, a "retry," a manual edit of the AI's output — all of these are signal, and most teams throw them away.

If you can't reconstruct the exact behavior of a single production request from your observability stack, you are flying blind.

Design for partial failure, not for "the model was wrong"

Deterministic systems fail loudly: a service is down, a database is unreachable, an exception is thrown. Non-deterministic systems fail subtly: the answer is plausible, fluently written, and quietly wrong.

This means "did it fail?" is no longer a binary question your code can answer for itself. Some of the most important production patterns we use:

Structured outputs with schema validation. If the model returns JSON, validate it. If validation fails, you have a clean retry path. If it passes but the values are nonsense, that's where the next layer catches it.

Confidence-aware routing. Not every decision needs full autonomy. Categorize requests by stakes. Low-stakes: let the model act. Medium-stakes: ask the model to act but let a human approve. High-stakes: use the model to draft, never to decide.

from enum import Enum
from dataclasses import dataclass

class Action(Enum):
    AUTO_APPLY = "auto_apply"        # model acts, no human in the loop
    QUEUE_FOR_REVIEW = "review"      # human approves before commit
    DRAFT_ONLY = "draft"             # human decides, model only suggests

@dataclass
class RoutingDecision:
    action: Action
    reason: str

def route(prediction, *, stakes: str) -> RoutingDecision:
    # High stakes always require a human, regardless of confidence.
    if stakes == "high":
        return RoutingDecision(Action.DRAFT_ONLY, "high-stakes path")

    if prediction.confidence >= 0.90 and stakes == "low":
        return RoutingDecision(Action.AUTO_APPLY, "high confidence, low stakes")

    if prediction.confidence >= 0.70:
        return RoutingDecision(Action.QUEUE_FOR_REVIEW, "medium confidence")

    return RoutingDecision(Action.DRAFT_ONLY, "low confidence")
from enum import Enum
from dataclasses import dataclass

class Action(Enum):
    AUTO_APPLY = "auto_apply"        # model acts, no human in the loop
    QUEUE_FOR_REVIEW = "review"      # human approves before commit
    DRAFT_ONLY = "draft"             # human decides, model only suggests

@dataclass
class RoutingDecision:
    action: Action
    reason: str

def route(prediction, *, stakes: str) -> RoutingDecision:
    # High stakes always require a human, regardless of confidence.
    if stakes == "high":
        return RoutingDecision(Action.DRAFT_ONLY, "high-stakes path")

    if prediction.confidence >= 0.90 and stakes == "low":
        return RoutingDecision(Action.AUTO_APPLY, "high confidence, low stakes")

    if prediction.confidence >= 0.70:
        return RoutingDecision(Action.QUEUE_FOR_REVIEW, "medium confidence")

    return RoutingDecision(Action.DRAFT_ONLY, "low confidence")

Two things worth flagging here. First, "confidence" is whatever signal you actually trust — logprobs, a calibrated classifier, an LLM-as-judge score, agreement across multiple samples. The score is only as useful as its calibration, so measure it. Second, the thresholds are product decisions, not engineering ones. They belong in config, reviewed by whoever owns the business impact of a wrong call.

Graceful degradation. When the model is slow, expensive, or refusing, what's the fallback? Cached result? Simpler model? Human escalation? "Sorry, try again"? Designing this is product work, not just infra work.

Cost is a feature

Compute is no longer free or even close to free, and unlike most cloud spend, AI cost scales linearly with usage. A feature that delights one user at three cents per call delights a thousand users at thirty dollars per hour, all day.

This needs to live in your engineering reviews from day one. Not as a side concern, but as a first-class design constraint, in the same column as latency and accuracy. Some questions worth asking before any feature ships:

  • What's the cost per successful outcome (not per call)?

  • What's the cost per failed outcome, including retries?

  • Where is the most expensive token going, and is it necessary?

  • Can a smaller model do 80% of the work and only escalate when needed?

We've seen teams cut their inference bill by an order of magnitude not by switching providers, but by tightening their prompts, caching aggressively, and being honest about which steps actually need a frontier model.

Humans are part of the architecture

The most reliable AI systems in production right now aren't fully autonomous. They have humans in the loop — sometimes invisibly, sometimes overtly — because the cost of being wrong outweighs the cost of being slow.

This isn't a temporary scaffolding to be removed once the model gets better. It's a design choice. Once you accept it, a lot of things get easier: you can ship sooner, you can iterate based on what humans actually correct, and you build a labeled dataset for the next round of evals as a side effect of normal operation.

The engineering work is making the human's job ergonomic. A review queue with the right context, the right diff, and the right keyboard shortcuts is a real product surface, not an afterthought.

What engineering teams should take from this

A few things that are easy to say and harder to internalize:

The deployment isn't the finish line — it's the start of the measurement. With deterministic software, shipping a feature mostly ends the engineering work. With applied AI, shipping is when you finally get to see how the system actually behaves at scale, against real inputs, with real users doing things you didn't anticipate.

Versioning matters more, not less. Prompts, eval sets, model versions, retrieval indices — all of these need to be versioned with the same rigor as your code. A silent prompt change is the new silent migration.

Hiring and team shape need to evolve. Pure backend engineers will struggle without exposure to evals and prompt engineering. Pure ML researchers will struggle with production reliability. The teams that ship well tend to have people who can hold both contexts.

Skepticism is a senior skill. The most valuable thing a senior engineer can bring to an applied AI project is the instinct to ask, of any impressive-looking output, "yes, but how often does it look like this, and what does it look like when it's wrong?"

Applied AI in production is not the same job as building deterministic software, and pretending it is — or that the model will "just get better" — is how you ship something that works in the demo and embarrasses you in the customer's hands.

Treat the non-determinism as the design problem, not the obstacle. The teams that do this are the ones whose AI features stop being a demo and start being a product.

No heading elements found. Showing placeholder content.