From PoC to Production

    The engineering discipline behind enterprise AI

    Anirban Ghatak
    Anirban Ghatak
    Oct 1, 2026
    The PoC trap: proving possibility is not proving production readiness in enterprise

    A successful AI proof of concept proves that a model, prompt or retrieval pattern can work under controlled conditions. Production asks a much harder question: can the same system remain dependable when regulations,real users, changing data and enterprise controls enter the picture?

    The variables multiply quickly:

    1. Latency, concurrency and cost become operational constraints rather than demo metrics.
    2. Source data changes and retrieval behavior drifts.
    3. Prompts become unpredictable, permissions tighten and every action needs an audit trail.
    4. Failures must be diagnosable and recoverable, not simply reproduced in a notebook.

    The hard part is rarely getting the first useful answer from an LLM. The hard part is creating an operating model around that answer: governed data, measurable quality, controlled deployment, observable behavior and a disciplined feedback loop.


    Key idea: The production unit is not the model. It is the entire AI service — plus the evidence that tells us whether that service is safe, useful and improving.


    Architecture matters because production AI is a system, not a model

    A practical production architecture connects governed enterprise data to retrieval, reasoning, serving and application consumption. For example on Databricks, that typically means Delta table and Unity Catalog controls, an AI Search or retrieval layer, RAG or agent logic, Model Serving, MLflow tracing and evaluation, and runtime governance.

    Figure 1. A simplified production AI reference architecture: governed data, retrieval, AI logic, serving, observability, evaluation and runtime control.

    The architecture is deliberately less about choosing one model and more about making the system repeatable. If a response is wrong, the engineering team should be able to reconstruct what data was retrieved, which prompt and model version ran, which tools were invoked, how long each step took, which policy allowed the action and what evidence supports the final answer.

    A useful technical nuance is tool governance. MCP is an open standard for connecting agents to tools and resources. In the current Databricks pattern, Unity Catalog can govern executable tool definitions such as functions, while Unity Gateway can govern MCP servers and access to them. The architectural principle is simple: open tool protocols can be used, but tool execution should still sit behind enterprise identity, permissions, credentials and audit controls.

    Build RAG as a governed knowledge product

    RAG is often introduced as a prompt technique: retrieve a few chunks, add them to context and ask the model to answer. In production, that framing is too narrow. Retrieval quality becomes part of the release process, which means it needs owners, tests, service expectations and observable failure modes.

    The source data must be governed. Chunking rules, embedding configuration and index construction should be captured as versioned build artifacts. External embedding APIs can change or deprecate models over time, so production teams should pin the exact embedding model version whenever the provider supports it. Where stronger reproducibility and lifecycle control are required, a self-hosted open-source embedding model served inside the platform can reduce dependence on an external model lifecycle.

    Search quality must also be measurable. Filters and reranking need predictable behavior, and the application needs an explicit response for insufficient evidence rather than encouraging the model to fill gaps confidently. The engineering question therefore changes from “Does the answer look good?” to “Can we prove why this answer was produced, and can we detect when the evidence is no longer good enough?”

    Evaluation should become a release gate

    Traditional software has deterministic tests. AI systems add a probabilistic layer, so release decisions need another form of evidence. A production team should maintain representative evaluation sets and score new versions for the dimensions that matter to the business: groundedness, relevance, correctness, retrieval quality, tool accuracy, safety, latency and cost.

    Agents add another dimension because quality is not only the final answer; it is also the path taken to reach it. Multi-step agents can loop, repeat tool calls or take unnecessarily long trajectories. Evaluation should therefore include trajectory-level checks such as step efficiency, tool-call sequence, maximum-step or loop budgets, and whether the agent reached the goal without getting lost in recursive reasoning.

    The objective is not to chase a perfect score. It is to make change measurable. A new prompt, retriever, model or agent policy should be able to demonstrate that it improves the desired behavior without silently breaking something that already worked.

    Release principle: A model change is not ready because it looks better in a demo. It is ready when the system can show measurable improvement against agreed quality and operating thresholds.


    Observability turns AI behavior from a black box into an engineering trace

    Once deployed, traditional infrastructure metrics are necessary but insufficient. CPU, memory and endpoint uptime can tell us whether the service is alive; they cannot explain why an answer was unsupported, why a tool call failed or why one request suddenly became expensive.

    Figure 2. Tracing decomposes an AI request into inspectable spans across retrieval, prompt construction, model inference, tool execution and the final answer.

    One way is to use MLflow tracing. Traces and spans let teams inspect the path from user question to retrieved evidence, prompt, model invocation, tool call and answer. They expose retrieval scores, context, parameters, token usage, latency and exceptions. In operational terms, this shortens the distance between “the AI gave a bad answer” and a specific engineering diagnosis.

    Observability also creates the evidence needed for continuous improvement. Production failures can become new evaluation cases instead of remaining anecdotal support tickets.

    Governance belongs inside the runtime architecture

    Governance is often treated as a checklist performed after an AI application has been built. Production AI requires the opposite approach. Data permissions, model access, tool authorization, audit logs, guardrails, cost visibility and human approval points should be designed into the runtime path.

    For agentic systems, static role-based access alone can be too coarse. A useful production pattern is to combine baseline privileges with context-aware controls such as attribute-based access control (ABAC), governed tags and row or column policies. Where tools act on behalf of a user, run-as-user or user-proxy execution identities help preserve the caller’s authorization boundary rather than allowing the agent to inherit a broad service identity.

    The question is therefore not only “Who can access this table or model?” It is also “Which tool can the agent invoke, under whose identity, against which data, with what limits, and how will that action be audited?” The moment an agent can update a record, trigger a workflow or call an external service, runtime authorization becomes part of application design.

    Close the loop: production telemetry should improve the next release

    The most mature operating model is circular rather than linear. Teams build, evaluate, deploy, observe, govern and then use production evidence to improve the next version. A failed retrieval becomes a new test case. A slow tool call becomes an optimization target. Repeated unsupported answers reveal a knowledge gap. A looping agent becomes a trajectory test. A policy exception becomes a control-design input.

    The operating loop: Build → Evaluate → Deploy → Observe → Govern → Improve → Build again


    This is the real distinction between an AI PoC and a production AI product: the product has a mechanism for learning from its own operation. For enterprises, the strategic question is no longer simply whether an LLM can solve a use case.

    The more important question is whether the organization can operate that use case with enough evidence to trust it, enough control to scale it and enough feedback to improve it continuously.

    “Production AI is not a deployment event. It is a continuously evaluated engineering loop”.

    Moving AI from PoC to production is not primarily a deployment challenge; it is an operating-model challenge. The real value comes from combining governed data, measurable quality, observable execution and secure access with a continuous improvement loop. The teams that treat AI as a continuously evaluated product—not a one-time experiment—will be the ones that scale it reliably.

    Continue the journey

    This article distills the core ideas from the webinar “From PoC to Production: Deploying & Observing Enterprise AI”.

    For engineers building customer-facing, production-ready AI systems, explore the ADaSci Certified Forward Deployed Engineer (CFDE) program.


    anirban.ghatak

    Anirban Ghatak

    Anirban Ghatak is a seasoned AI & Data Science leader with over 21 years of experience building, scaling, and leading analytics, BI, and data science business units with full P&L ownership and C-level reporting. An alumnus of BITS Pilani and IIM Indore, Anirban is an ex-founder and intrapreneur specialized in taking enterprise AI practices to scale. He is a prominent industry voice and speaker on the evolution of work in an Agentic AI world, advocating for the transition from systems of execution to hybrid systems of autonomous orchestration and robust AI governance.

    Get Credentialed

    More Articles

    Comments (0)

    Join the conversation

    Sign in to comment

    Loading comments...