Baseline Microsoft Foundry Chat Reference Architecture in an Azure Landing Zone
A deep‑dive into the reference architecture that separates platform‑managed shared services from workload‑specific AI and chat components, enabling enterprises to maintain strict governance, security, and cost control while leveraging Microsoft Foundry for generative AI development.
Executive Introduction
Enterprises embarking on generative‑AI initiatives often grapple with a tangled web of infrastructure responsibilities. On one side, platform teams must provide a secure, governed, and cost‑efficient foundation—networking, identity, policy, monitoring, and shared services. On the other side, workload teams need a rapid, sandboxed environment to build, test, and deploy AI models, chat agents, and the associated data stores.
The Baseline Microsoft Foundry Chat Reference Architecture in an Azure Landing Zone offers a proven pattern that aligns these opposing needs. By treating the Foundry resource as a workload‑owned asset and delegating all shared infrastructure to a platform landing zone, organizations can:
- Achieve clear ownership and accountability through subscription democratization.
- Enforce consistent governance and policy across multiple business groups using Azure Policy and DeployIfNotExists (DINE) constraints.
- Maintain strict network segmentation with hub‑spoke topologies, private endpoints, and controlled egress via Azure Firewall.
- Reduce operational friction by centralizing DNS, DDoS protection, and bastion access in a connectivity subscription.
For enterprise IT leaders, this architecture delivers a repeatable blueprint that balances agility with compliance, minimizes the risk of drift, and provides a solid footing for scaling AI workloads across the organization. The following article unpacks the technical details, implementation considerations, and operational implications of this pattern, while highlighting why it matters to modern enterprise IT and how Escape Business Solutions (EBS) can help you operationalize it.
Technical Architecture Overview
The reference architecture is divided into two logical subscription groups: **application landing zone** (where workload resources reside) and **platform landing zone** (where shared, cross‑cutting services live). This separation mirrors Azure landing‑zone design principles and introduces a clear “who‑owns‑what” contract between platform and workload teams.
Application Landing Zone Subscription
- Spoke Virtual Network – Contains all workload components, including App Service, Application Gateway, private endpoints for PaaS services (Azure Storage, Key Vault, Foundry, Azure AI Search, Cosmos DB), and a dedicated subnet for agents.
- Subnet Design – The spoke includes:
snet-appGateway– Front‑door for ingress traffic.snet-buildAgents– Development and testing agents.snet-jumpBoxes– Administrative jump hosts.snet-appServicePlan– App Service plan with zone‑redundant instances.snet-agentsEgress– Egress path for chat agents, routed through a data proxy.snet-privateEndpoints– Hosts private endpoints for all dependent PaaS resources.
- Network Security Groups (NSGs) – Applied per subnet to enforce segmentation and limit lateral movement.
- User‑Defined Routes (UDRs) – Force all outbound traffic from the
snet-agentsEgressand other workload subnets through the hub’s Azure Firewall (or a connectivity‑subscription gateway). This provides a single point of inspection and logging. - Foundry Resource & Projects – The workload team owns a single Foundry resource (AI application platform) with projects that expose chat agent endpoints. Each project hosts the Agent Service runtime, content‑safety logic, and model‑as‑a‑service (MaaS) deployments.
- Agent Service & Dependencies – Agent Service runs in its own dedicated subnet, communicating via a single‑tenant data proxy. It stores state, chat history, and file storage in Azure Cosmos DB (NoSQL), Azure Storage, and Azure AI Search, respectively.
- Web Front‑End – Azure App Service (zone‑redundant) hosts the chat UI and API apps. The application code is packaged as a ZIP file stored in an Azure Storage account and mounted at runtime.
- Application Gateway & WAF – Acts as a reverse proxy, terminating TLS from external users and routing traffic to App Service. The integrated web‑application firewall (WAF) blocks common attacks.
- Key Vault – Securely stores the Application Gateway’s TLS certificate and any other secrets required by workload components.
- Observability – Azure Monitor, Azure Monitor Logs, and Application Insights are configured to ingest logs and metrics from all workload resources.
- Governance – Azure Policy and DINE policies are applied directly to this subscription (or via a management group) to enforce naming conventions, tag requirements, and resource quotas.
Platform Landing Zone Subscription
- Hub Virtual Network – Centralized networking hub containing Azure Firewall, Azure Bastion, and optional VPN Gateway/ExpressRoute.
- Azure Firewall – Manages all egress traffic from the spoke (including agent traffic) and can be extended to inspect inbound traffic if required by governance policies.
- Azure Bastion – Provides jump‑box access to VMs in the spoke without exposing RDP/SSH ports to the internet.
- Connectivity Resources – DNS Private Resolver, Private DNS zones for Private Link, DDoS Protection plan for the Application Gateway’s public IP, and the underlying ExpressRoute/VPN gateways.
- Management Groups & Policies – Platform‑team‑controlled management groups enforce cross‑subscription governance, including DINE policies that may constrain resource placement in the application landing zone.
- Network‑based Services – Centralized DNS resolution for the spoke, cross‑premises connectivity, and network virtual appliances (NVAs) for advanced filtering.
Key Ownership Contracts
The architecture codifies ownership through subscription democratization: each subscription is a “vended” resource, with clear responsibilities documented in a checklist that both platform and workload teams must sign off on. The platform team owns all resources residing in the connectivity subscription (hub networking, firewall, bastion, DNS, DDoS protection). The workload team owns everything deployed within the application landing zone subscription (including the Foundry resource, agent services, data stores, and front‑end web components). This delineation eliminates ambiguity and simplifies cost‑allocation reporting.
How the Technology Works
Understanding the mechanics of each component helps teams design, troubleshoot, and optimize the architecture for their specific workloads.
Foundry Resource & Projects
- Foundry as a Platform – Provides a unified SDK, model management, and agent orchestration. In this pattern, the workload team creates a single Foundry resource (e.g.,
/subscriptions/<app‑sub>/resourceGroups/foundry-rg/providers/Microsoft.CognitiveServices/accounts/foundry‑ai). Within this resource, projects act as logical containers for different chat capabilities (e.g.,project‑customer‑support,project‑internal‑hr). - Agent Service Runtime – Part of the Foundry project, Agent Service is a cloud‑native runtime that hosts prompt agents (or containerized hosted agents). It reads configuration from Key Vault, authenticates via managed identities, and forwards user queries through the data proxy in the
snet-agentsEgresssubnet. - Model‑as‑a‑Service (MaaS) – Generative AI models (e.g., GPT‑4, Llama) are deployed as MaaS endpoints within the Foundry resource. These endpoints are secured with content‑safety filters that can be configured per project.
Agent Service Data Path
Agent traffic follows a tightly controlled path:
- User sends a chat request via the App Service front‑end (HTTPS terminated at Application Gateway).
- App Service forwards the request to Agent Service (internal DNS name resolved via hub DNS).
- Agent Service initiates outbound traffic from the
snet-agentsEgresssubnet through the single‑tenant data proxy. - Data proxy uses UDR to send traffic to Azure Firewall for inspection and logging before reaching the internet or downstream services (e.g., external APIs, grounding context stores).
- Inbound responses (including model outputs) travel back through the same path and are relayed to the user.
Knowledge Store & Retrieval Augmented Generation (RAG)
- Foundry IQ powered by Azure AI Search serves as the workload’s knowledge base. Prompts are first processed by a lightweight query transformer, then submitted to Azure AI Search to retrieve relevant document vectors.
- Search results are passed as grounding data to the generative model endpoint, ensuring responses are factually anchored to enterprise‑specific content.
- The indexing pipeline (triggered by data‑scientists) runs as an Azure Function hosted in the application landing zone and writes indexed data to Azure AI Search and optionally to Cosmos DB for low‑latency access.
Security & Identity Flow
- Managed Identities – Each workload component (App Service, Agent Service, Functions) is assigned a system‑assigned or user‑assigned managed identity, granting access to Azure Key Vault, storage accounts, and Foundry resources without embedding secrets.
- Private Endpoints – All PaaS dependencies (Storage, Key Vault, Foundry, AI Search, Cosmos DB) expose private endpoints within the spoke. This eliminates public‑internet exposure and ensures traffic never leaves the Microsoft backbone.
- Conditional Access – Azure AD conditional access policies can be applied to the App Service (via Azure AD App Registration) to enforce multi‑factor authentication for the chat UI.
Observability & Governance
- Azure Monitor + Application Insights capture request latency, error rates, and custom telemetry from Agent Service and the RAG pipeline.
- Azure Policy enforces tagging, resource naming conventions, and resource limits (e.g., maximum number of Foundry projects per resource group).
- DeployIfNotExists (DINE) policies guarantee that required network configurations (NSG rules, UDR entries) are present whenever new resources are added.
- Change Tracking – Azure Policy’s audit mode and Log Analytics workspaces provide a centralized view of compliance status across both landing zones.
Implementation Considerations
Deploying this architecture requires careful coordination, clear documentation, and a disciplined approach to resource organization.
Prerequisites & Stakeholder Alignment
- Established Azure landing‑zone framework (hub‑spoke topology, connectivity subscription, management groups).
- Platform team has already provisioned Azure Firewall, Bastion, DNS Private Resolver, and DDoS Protection in the connectivity subscription.
- Workload team has identified business criticality for each chat use‑case to inform management‑group placement.
- Both teams have completed the subscription‑vending questionnaire (see below) to capture required IP address spaces, VM families, and policy constraints.
Network Design & IP Allocation
- Spoke Address Space – The platform team allocates a /16 (or smaller) to the spoke based on the workload’s subnet list. Each subnet must be /24 or larger to accommodate current and future resources.
- DNS Resolution – Hub‑based DNS (Azure Private DNS zones) resolves private endpoint names. The workload team should not configure custom DNS settings on the spoke; relying on hub DNS guarantees consistency.
- Egress Control – UDR entries in the spoke force traffic from
snet-agentsEgress(and optionally other subnets) to the hub’s Azure Firewall. This ensures all outbound flows are logged and can be throttled. - Inbound Termination – Application Gateway is placed in the spoke, with its public IP protected by the DDoS Protection plan. If governance mandates an additional inbound inspection layer, the platform team can enable Azure Firewall in front of Application Gateway (requires extra UDRs and firewall rules).
Resource Organization & Foundry Ownership
- Single Foundry Resource – The workload team provisions one Foundry account and delegates projects to different business units or product teams. This centralizes cost tracking and licensing while still allowing per‑project governance.
- Project Isolation – Each project resides in its own resource group, with distinct keys, content‑safety configurations, and model deployments. Role‑based access control (RBAC) at the project level ensures developers only see their own artifacts.
- Data Store Ownership – Cosmos DB, Azure Storage, and Azure AI Search instances are tagged with a “workload‑owner” tag and are exclusively managed by Agent Service. No cross‑workload sharing is allowed to maintain data‑privacy boundaries.
Governance via Azure Policy & DINE
- Platform‑Level Policies – Policies such as
AllowedResourceLocations,AllowedSkus, andTagRGare applied at the management group level covering both platform and application landing zones. - Workload‑Specific Policies – Within the application landing zone subscription, policies may enforce a maximum number of Foundry projects, required tags for model assets, and mandatory use of private endpoints for all PaaS resources.
- DINE Enforcement – Guarantees that when a new subnet is created, the corresponding NSG rules and UDR entries are automatically provisioned to maintain network security baselines.
Subscription Vending Checklist
The platform team typically follows a structured questionnaire to capture the necessary inputs before provisioning the application landing zone subscription. An example checklist includes:
| Category | Question |
|---|---|
| Networking | What is the desired address space for the spoke network and the required subnet sizes for each workload component? |
| Security | Do we require inbound inspection via Azure Firewall or any custom NSG rules beyond the baseline? |
| Identity | What Azure AD applications and managed identities are needed for App Service, Agent Service, and functions? |
| Monitoring | Which Log Analytics workspaces or Azure Monitor configurations should be linked to this subscription? |
| Governance | What tags, naming conventions, and Azure Policy assignments must be applied at subscription level? |
| Cost Management | Are there any cost‑allocation tags or budget alerts required for this workload? |
Both teams review the checklist, negotiate any constraints, and then the platform team “vend” the subscription by applying the agreed‑upon configuration (network peering, NSG baseline, policy assignments). The workload team then proceeds with deploying Foundry, agents, and chat UI within the approved parameters.
Security and Governance
Security is woven into every layer of the reference architecture, from network segmentation to identity management.
Network Security Controls
- Spoke NSGs – Define allow/deny rules per subnet (e.g., allow HTTPS from internet to
snet-appGateway, allow traffic fromsnet-appServicePlantosnet-privateEndpoints, deny all else). - Hub Azure Firewall – Centralized egress inspection, optional inbound proxy, and application‑level logging. Rules are designed to allow only required outbound destinations (e.g., internet, cross‑premises, other landing zones).
- Private Endpoints – All PaaS services are accessed via private endpoints, removing public IP exposure. DNS resolution for these endpoints is handled by Azure Private DNS zones managed by the platform team.
- DDoS Protection – Applied to the public IP of Application Gateway, providing automatic mitigation against volumetric attacks.
Identity & Access Management
- Managed Identities – Reduce credential leakage by using Azure AD–managed identities for Azure resources.
- Key Vault Access Policies – Limited to specific managed identities; no developer ever handles a secret directly.
- Azure AD Conditional Access – Enforce multi‑factor authentication for the chat UI, and restrict access based on device compliance.
- Role‑Based Access Control (RBAC) – Granular permissions at subscription, resource group, and resource level. Platform team retains ownership of networking resources, while workload team controls Foundry, agent, and data‑store resources.
Policy‑Based Governance
- Azure Policy – Enforces naming conventions, resource tags, and location constraints. Example policy:
LogonHourscan be used to limit when administrative VMs in the jumpbox subnet may run. - DINE Policies – Ensure that any new subnet automatically gets a baseline NSG rule (e.g., allow traffic to private endpoints) and a UDR pointing to the hub firewall.
- Audit & Compliance – Azure Policy’s audit mode provides a snapshot of non‑compliant resources, enabling remediation before production rollout.
Operational Security Considerations
- Segregation of Duties – Platform and workload teams should maintain separate Azure AD groups to enforce principle of least privilege.
- Change Management – All changes to networking (UDR updates, firewall rules) must follow the platform team’s change‑request process.
- Backup & Disaster Recovery – While not part of the baseline, Cosmos DB and Storage accounts should have backup policies configured (e.g., point‑in‑time restore for Cosmos DB, geo‑redundant storage for blob data).
- Monitoring & Alerting
- Critical alerts include: Azure Firewall denial of traffic, excessive error rates from Application Insights, and policy violations logged in Log Analytics.
Operational Implications
Running a chat workload at enterprise scale introduces operational complexities that go beyond pure deployment.
Observability & Troubleshooting
- Centralized Logs – Azure Monitor aggregates logs from Application Gateway (WAF logs), App Service (App Service logs), Agent Service (custom telemetry), and the RAG pipeline (Function logs).
- Dashboarding – Azure Portal dashboards can surface key metrics: active chat sessions, model latency, search index freshness, and firewall throughput.
- Root Cause Analysis – When an agent fails to reach a grounding context store, combine firewall logs, DNS query logs, and Agent Service telemetry to pinpoint whether the issue is network, authentication, or data store related.
Scaling & Performance
- App Service Scaling – Auto‑scale based on CPU or request count; the architecture recommends a minimum of three instances across availability zones for high availability.
- Agent Service Concurrency – Agent Service instances can be scaled horizontally; each instance runs in its own subnet (multiple
snet-agentsEgresssubnets) to distribute load. - AI Search Indexing
– Search indexes can be sharded; the platform team should monitor index size and refresh cadence to avoid performance degradation.
- Network Bandwidth – Ensure the hub’s ExpressRoute/VPN circuit has sufficient bandwidth to handle peak agent traffic, especially if large documents are processed.
Cost Management
- Resource Tagging – Enforce a “CostCenter” tag on all workload resources to enable Azure Cost Management reporting.
- Reserved Instances – Consider Azure Hybrid Benefit or Reserved Instances for predictable compute (App Service plans, Cosmos DB throughput).
- Monitoring Spend
- Azure Monitor and Log Analytics usage should be tracked; set budget alerts for unexpected spikes.
Change & Release Management
- CI/CD Pipelines
- Use Azure DevOps or GitHub Actions to automate deployment of web UI, functions, and Agent Service configurations. Ensure pipelines run against non‑production environments first.
- Versioned Infrastructure
- Store ARM templates, Bicep files, and policy definitions in a private Git repository with branch protection to maintain auditability.
Common Pitfalls & How to Avoid Them
- Incorrect DNS Resolution – Forgetting to configure hub DNS for private endpoints can cause agent connectivity failures. Mitigation: Verify DNS Private Resolver rules and test resolution from a workload VM.
- Missing UDR Entries – Agents bypassing the firewall can expose internal traffic. Mitigation: enforce UDR at the subnet level and validate traffic flow with Network Watcher.
- Over‑privileged Managed Identities
- Granting a managed identity rights to all storage accounts can break isolation. Mitigation: Apply principle of least privilege; use conditional access for admin tasks.
- Policy Drift – Adding resources outside the approved subscription can break governance. Mitigation: Use Azure Policy’s Deny or DoNotUse effects and enforce strict RBAC.
- Neglecting Backup – Assuming Cosmos DB or Storage is immutable leads to data loss. Mitigation: Enable point‑in‑time backup for Cosmos DB and geo‑redundant storage for blobs.
- Assuming Workload Owns Everything
- Attempting to manage networking in the application landing zone without platform team cooperation can cause connectivity issues. Mitigation: Follow the subscription‑vending checklist and let the platform team own hub resources.
Why this matters to enterprise IT
The baseline Microsoft Foundry Chat Reference Architecture aligns directly with core enterprise IT objectives: **governance, security, cost control, and scalability**.
- Clear Governance – By codifying ownership through subscription democratization, enterprises avoid the “everybody‑owns‑everything” chaos that often leads to compliance breaches.
- Enhanced Security Posture – Private endpoints, centralized firewall inspection, and managed identities collectively reduce the attack surface dramatically compared to publicly exposed AI services.
- Predictable Cost
- Centralized cost‑allocation tags and reserved‑instance planning enable accurate budgeting for AI initiatives.
- Scalable Architecture
- Horizontal scaling of App Service, Agent Service, and AI Search ensures the system can support increasing user loads without redesign.
- Operational Efficiency
- Shared platform services (DNS, Bastion, DDoS protection) are provisioned once and reused across multiple workload teams, reducing operational overhead.
- Compliance & Auditability
- All changes flow through Azure Policy and DINE, providing a tamper‑evident trail for auditors and regulators.
Consequently, enterprise IT can confidently adopt generative‑AI chat capabilities, knowing they are built on a foundation that meets internal security policies, external regulatory requirements, and financial accountability.
EBS consulting perspective
From an Escape Business Solutions consulting standpoint, this reference architecture provides a ready‑made framework for accelerating AI‑driven product delivery while upholding enterprise‑grade controls. Our methodology emphasizes:
- Assessment & Blueprinting – We begin with a discovery workshop to capture existing landing‑zone components, identify gaps, and tailor the baseline to your specific regulatory context.
- Design & Governance Alignment – Leveraging Azure Well‑Architected Framework principles, we design the platform and workload subscriptions, defining RBAC, policy, and tagging strategies that align with your cost‑center and business unit structures.
- Implementation & Validation – Our engineers provision the hub network, firewall, DNS, and bastion services, then guide the workload team through Foundry resource creation, agent onboarding, and chat UI deployment. We run end‑to‑end integration tests to verify connectivity, security controls, and performance baselines.
- Enablement & Knowledge Transfer – We deliver comprehensive runbooks, monitoring dashboards, and a training session for platform and workload operators, ensuring internal teams can maintain and evolve the environment.
- Continuous Improvement
- Through ongoing managed services, we monitor policy compliance, cost trends, and security incidents, recommending iterative refinements to keep the AI platform aligned with business goals.
By partnering with EBS, organizations can reduce time‑to‑value for generative‑AI projects, avoid common pitfalls, and achieve a hardened, governable cloud environment that scales with future AI ambitions.
Practical next steps
If your organization is evaluating or currently building a generative‑AI chat capability, consider the following actionable roadmap:
- Stakeholder Alignment – Convene platform and workload leads to review the subscription‑vending checklist and agree on networking, identity, and governance requirements.
- Baseline Landing‑Zone Setup – Ensure the platform landing zone includes Azure Firewall, Bastion, DNS Private Resolver, and DDoS Protection. Document these as “already provisioned” dependencies.
- Resource Allocation – Request the required IP address space for the application landing zone spoke from the platform team. Confirm subnet sizes for each component listed in the architecture.
- Policy Definition – Draft Azure Policy and DINE definitions for the workload subscription (naming, tagging, private endpoint enforcement). Validate them in a non‑production subscription.
- Foundry Provision – Create a single Foundry resource and initial projects under the workload subscription. Assign appropriate RBAC roles (e.g., Cognitive Services Contributor) to the development team.
- Agent Service Onboarding
- Deploy Agent Service instances in the
snet-agentsEgresssubnet, configure the data proxy, and link the associated Cosmos DB, Storage, and AI Search accounts.
- Front‑End Deployment
- Package the chat UI as a ZIP, deploy it to Azure Storage, mount it in an App Service plan with zone‑redundant instances, and configure Application Gateway with WAF.
- Connectivity Validation
- Run network tests (Azure Network Watcher, DNS resolution checks) from representative VMs to confirm agent egress flows through the firewall and private endpoints resolve correctly.
- Observability Setup
- Configure Azure Monitor, Application Insights, and Log Analytics workspaces to capture logs from all components. Build basic alerts for high error rates or firewall denials.
- Documentation & Handover
- Create an operational runbook that includes change procedures, troubleshooting steps, and escalation contacts. Conduct a formal handoff to the platform and workload operation teams.
Completing these steps provides a solid foundation for scaling the chat solution across multiple business units while preserving the governance and security standards demanded by enterprise IT.
Conclusion
The Baseline Microsoft Foundry Chat Reference Architecture in an Azure Landing Zone encapsulates a disciplined approach to delivering generative‑AI capabilities at enterprise scale. By clearly separating platform responsibilities (networking, security, governance) from workload ownership (Foundry, agents, data stores), organizations gain visibility, control, and the ability to enforce consistent policies across diverse business groups.
For enterprises seeking to accelerate AI adoption without compromising on compliance or cost predictability, this architecture offers a pragmatic, battle‑tested pattern. Escape Business Solutions can guide you through every phase—from assessing your current environment and designing a tailored landing‑zone blueprint to implementing the solution and establishing ongoing operational excellence.
Ready to modernize your AI platform while maintaining enterprise‑grade governance? Contact EBS today to begin a guided workshop that aligns your business objectives with a secure, scalable Azure landing‑zone architecture.
EBS Consulting Advice
If your organization is evaluating Baseline Microsoft Foundry Chat Reference Architecture in an Azure Landing Zone – Azure Architecture Center, do not treat the technology decision in isolation. Start with the business outcome, current architecture, security and identity controls, operational constraints, migration dependencies and governance requirements. A practical assessment should identify the current-state gaps, prioritize the risks and define an implementation roadmap with measurable outcomes.
EBS can help assess the environment, develop the architecture and modernization roadmap, and translate the technical options into an actionable business plan. Relevant EBS services: Microsoft Azure consulting Escape Cloud Microsoft Solution Assessments.
Have a technology challenge? Email info@escapebusinesssolutions.com to describe your situation. We welcome questions, consulting discussions and requests for a proposal.
Discover more from Escape Business Solutions
Subscribe to get the latest posts sent to your email.
