Applied AI
Building AI Systems That Actually Reach Production
What changes when an AI prototype becomes a production system: workers, infrastructure, reliability, observability, and controlled execution.
· 12 min read
AI prototypes are easy to make impressive. A useful prompt, a model call, and a convincing response can demonstrate that an idea is worth pursuing.
Production changes the question. It is no longer only whether the model produces something useful. It is whether the product can account for the work when execution is delayed, repeated, interrupted, or only partially successful. The problem becomes a distributed systems problem as well as an AI problem.
The lessons here come from a product backend that combines AI generation and drafting with asynchronous processing and external business systems. Its implementation and regression tests provide the evidence; the diagrams and examples below deliberately generalize the domain. They describe architectural mechanisms, not an audit of every deployed behavior.
The prototype is not the product
A prototype can treat a model response as the end of the operation. A product has to decide what that response means for an existing workflow.
In the system behind this article, generated output sits inside application processes with persistent records, execution states, and external synchronization. There are also non-AI workloads: importing data, delivering notifications, and reconciling information with other systems. They compete for resources and have their own failure modes.
This changes the definition of completion. A successful provider response does not necessarily mean the result was saved, the relevant workflow advanced, or an external system accepted an update. Those are separate events with separate owners.
The product needs to distinguish accepted work from finished work, a recoverable interruption from a terminal failure, and an already-completed operation from a new request. Without those distinctions, an impressive model response can coexist with a broken user experience.
That is the first architectural shift: the unit of reasoning becomes the product operation, not the model invocation.
The LLM becomes one component
The backend contains LangGraph-based drafting alongside other AI generation integrations. That is evidence of applied AI, but it is not evidence that every task is an autonomous agent—or that the entire application belongs inside an agent graph.
Application services still coordinate persistence, validate inputs, and decide which operation should run. Workers execute queued tasks. Integration gateways handle external calls. Database records capture progress and results. Redis supports queueing and coordination.
These boundaries matter because the components fail differently. A model call can time out. A worker can disappear. A transaction can roll back. An external service can reject an update. Combining them into one opaque “AI workflow” hides the decisions needed to recover.
- User
- LLM
- Response
- Product / API
- Queue
- Worker
- AI + integrations
- Persisted result
The diagram is intentionally simplified. It does not reproduce an internal deployment topology or imply that every operation follows an identical route. Its point is that model execution belongs within a larger application boundary.
The engineering question becomes: what must be true before the model runs, what can happen while it is running, and what must be committed before the product reports completion?
Asynchronous work changes the failure model
The implementation uses Celery with Redis-backed task delivery and results. It routes different classes of work to different queues, including dedicated queues for longer or heavier processing.
That separation is more consequential than adding a background-task decorator. A long-running operation should not necessarily occupy the same execution path as a short request or periodic housekeeping. Separate queues make workload boundaries explicit, although queue separation alone does not prove that resources are isolated in every deployment.
Moving execution away from the HTTP lifecycle also changes the contract. The request can initiate work, but the worker’s eventual result needs an identity and a place to live. The user-facing application cannot depend exclusively on the original request still being open.
The code represents that distinction through execution records and persistent statuses. Some tasks retry with delays and backoff. When retries are exhausted, task wrappers can record an error so the broader workflow can resolve instead of remaining indefinitely in progress. Recovery logic also looks for work that has become stuck.
- Record intent
- Enqueue
- Execute
- Persist outcome
Result + terminal state
Transient failure + delay
Recorded error + recovery decision
These mechanisms address different questions. A queue answers where work waits. A retry policy answers whether an attempt should run again. An execution record answers what the product currently knows. Recovery code answers what to do when those views stop agreeing.
It is tempting to collapse all four into “the job status.” Keeping them conceptually separate makes failures easier to diagnose and prevents transport-level success from being mistaken for product-level completion.
Reliability becomes a product feature
One particularly useful lesson appears in the broker configuration and a dedicated regression test: acknowledgment timing and execution limits must agree.
With late acknowledgment, a task is acknowledged after execution rather than before it. Worker-loss handling can then allow unfinished work to be delivered again. But a Redis broker also has a visibility window. If that window expires while an operation is still running, another worker can receive the same message.
The repository explicitly tests that the visibility timeout exceeds the configured task time limits. The lesson is the relationship between those settings, not a particular timeout value. A setting that looks like infrastructure plumbing can determine whether the product performs overlapping work.
Late acknowledgment is therefore not a complete reliability strategy. It trades one risk—work disappearing after early acknowledgment—for another: repeated execution. The task and its persistence boundary must be designed to tolerate that possibility.
The AI gateway also distinguishes failure classes. Rate limits can honor a provider’s retry guidance; server and network failures can use backoff. Other client errors fail rather than being retried indiscriminately. A separate external-system consumer treats authentication failures differently from rate limits and server failures.
Those choices express product judgment. Repeating a request cannot fix invalid credentials. Retrying immediately during throttling can increase pressure. A transient transport failure may be worth another attempt, but only within the operation’s time budget.
Retries exist at more than one layer here: provider requests and worker tasks. That makes their interaction worth examining. A task retry can repeat earlier provider attempts; retry counts should not be interpreted independently of the total work they can produce. The inspected code does not justify a claim that every path shares one global retry budget.
State is more than model context
Model context describes the information used to produce an answer. Application state describes what the system is allowed to do next and what has already happened. They should not be confused.
A drafting graph can carry the inputs needed to generate text. That state does not, by itself, establish whether a workflow has been approved, whether its result has been stored, or whether an external update succeeded.
The backend uses SQL-backed records for application and execution state. Redis serves other roles, including task transport, results, and coordination. An expiring broker result is not a substitute for a durable product record: its lifetime and purpose are different.
The generation path provides a concrete example. It checks an execution record before doing expensive work, ignores terminal states, and uses a guarded transition when persisting success. The result, completion transition, and aggregate workflow update are committed together in the database.
That transaction protects the relationship between “completed” and “a result exists.” It does not make the remote generation call part of the database transaction.
- Check state
- External / AI call
- Check current state
Persist result + mark completion + update aggregate state
The implementation releases database locks before long external calls and reacquires protection around persistence. Holding a row lock across unpredictable network latency would turn an external dependency into a database contention problem.
The trade-off is important: another attempt may still spend computation before discovering that completion has already been recorded. Protecting the persisted outcome is different from guaranteeing that a provider was invoked exactly once.
This is why I would describe the mechanism as guarded, retry-aware persistence rather than claim end-to-end exactly-once execution. The distinction matters for cost, external side effects, and honest explanations of reliability.
External effects need their own boundary
An application database and an external business system do not share a local transaction. Saving an internal change and calling an external API creates a partial-failure boundary, even when neither operation involves AI.
The backend includes transactional outbox records, periodic dispatch, and consumers that track pending work, processing, completion, retry information, and errors. Tests cover cases such as skipping completed events and recording permanent failures.
An outbox gives intended delivery a durable representation. Instead of relying solely on a request to perform an external update immediately, later processing can find work that still needs attention. This is an application concern, not a property the LLM can provide.
The consumer also needs to understand outcomes. A confirmed external update can move the record to completion. A retryable failure should preserve enough information for another attempt. A permanent rejection needs a terminal state and a useful reason.
However, an outbox does not remove uncertainty at the remote boundary. If a remote operation succeeds and local completion is not recorded, a later attempt may repeat the call. Local “already completed” checks cannot prove what happened remotely before a crash.
The repository contains several defenses against repeated work, including webhook deduplication and guarded execution transitions. They solve specific boundaries. They should not be generalized into a promise that every external operation is idempotent. Reliability comes from identifying where each protection applies—and where it does not.
Infrastructure becomes part of AI engineering
The infrastructure code describes containerized services on AWS ECS/Fargate, managed database and Redis resources, task roles, secrets injection, and service-level resource configuration. Deployment workflows build container images and update ECS task definitions and services. There are also test workflows and migration checks.
None of that improves a prompt directly. All of it changes whether the product can operate the prompt-dependent workflow.
Worker memory and concurrency affect what can run together. A container replacement can interrupt execution. Network boundaries affect whether the database and external integrations are reachable. Secret configuration affects whether a service can authenticate. Schema changes affect whether a running worker can read and write its records correctly.
These are reasons to treat deployment and recovery as part of the application design. For example, configuring task redelivery only helps if the task’s persistence logic can handle another attempt. A database migration check is valuable because application code and durable state evolve together.
The presence of Terraform and CI workflows demonstrates those mechanisms in code; it does not establish that every environment has the same availability properties, that every deployment is zero-downtime, or that every release gate is enforced. I would want runtime and operational evidence before making those claims.
For an AI Product Engineer, the useful scope is broader than the provider SDK: understanding how a model-dependent feature behaves through a release, an interruption, and a recovery path.
Observability should answer operational questions
The repository contains structured logging, Prometheus-style metrics, and infrastructure for operational dashboards and log aggregation. Metrics include processing volume, duration, errors, last-success timestamps, and failed-work backlog size. The AI integration layer also configures LangSmith tracing, including initialization in worker processes.
These signals answer different questions. A model trace helps inspect an invocation. A worker log helps explain an attempt. A backlog metric helps identify accumulation. A last-success timestamp can reveal a process that has quietly stopped completing useful work.
The product operation needs to be recognizable across those views. Execution and event identifiers in structured logs are a useful basis for that investigation; public examples do not need to reveal the actual identifier formats or internal event names.
Still, instrumentation is not the same as complete observability. Tracing is configurable, and the inspected files do not establish comprehensive model-quality evaluation, complete cross-service correlation, or a measured reliability objective.
The practical lesson is to design around questions: which operation failed, where did it stop, can it be retried safely, and what state will the user see? A dashboard that cannot connect its signals to those questions may show activity without explaining the product’s behavior.
Controlled execution belongs to the application
Control does not have to mean an elaborate agent-safety framework. Often it means keeping ordinary application rules authoritative when AI enters the workflow.
The backend has authorization dependencies on protected routes and explicit review transitions. Some transitions are blocked when unresolved errors exist, and conflicting state changes can produce a conflict response. The generation integration also uses structured model output and validates application inputs rather than treating arbitrary text as executable instructions.
These are distinct controls. Authorization decides who can request an operation. State checks decide whether it is currently valid. Schema validation constrains what the application accepts. Review transitions determine when a process may advance.
A well-formed model response does not grant permission to perform an action. Equally, a prompt instruction is not a replacement for an application-level check. The deterministic boundary remains important even when the model is useful within it.
The inspected implementation supports describing controlled workflows, not a universal claim that every AI output is reviewed or that every possible unsafe action is prevented. That narrower description is both more accurate and more useful for engineering decisions.
What I learned
The strongest shift is from evaluating a response to accounting for an operation. The model is one participant in a product whose execution crosses process, persistence, and service boundaries.
Several lessons follow from the mechanisms above:
- Asynchronous execution introduces responsibilities, not just capacity. Accepted work needs durable identity, meaningful states, and a recovery decision when execution stops progressing.
- Retries and idempotency must be discussed together. Redelivery can preserve work while creating duplicate attempts. Protecting a stored result is not the same as preventing every repeated external call.
- Infrastructure settings shape product behavior. Broker visibility, task limits, concurrency, and container lifecycle are application concerns when they determine how work completes.
- Observability needs an operation to follow. Model traces, logs, and metrics are complementary; none alone proves that the user-facing workflow is healthy.
- AI does not replace product boundaries. Permissions, validation, state transitions, and review decisions still determine what the system may do.
Building production AI systems therefore means carrying the product problem through the whole execution path. Better model output matters. So does knowing what happens when the worker disappears, the integration refuses a request, or the database has not recorded the result yet.
That is where applied AI, backend engineering, cloud infrastructure, and product thinking become the same engineering problem.