Skip to main content Get to know Azure Microsoft as Customer Zero View all products (200+) Microsoft Foundry Azure Copilot GitHub Copilot Azure Kubernetes Service (AKS) Azure Cosmos DB Azure Database for PostgreSQL Azure Arc Microsoft Fabric Linux virtual machines in Azure Foundry Models Foundry Agent Service Foundry IQ Foundry Tools Foundry Control Plane Observability in Foundry Control Plane Azure OpenAI in Foundry Models Azure Speech in Foundry Tools Azure Machine Learning View all databases Azure Cosmos DB Azure DocumentDB Azure SQL Azure Database for PostgreSQL Azure Managed Redis Microsoft Fabric Azure Databricks Linux virtual machines in Azure Windows Server on Azure Azure Functions Azure Virtual Machine Scale Sets Azure API Management Azure Container Apps Azure Kubernetes Service (AKS) Azure Kubernetes Fleet Manager Azure Container Registry Azure Red Hat OpenShift Azure Container Instances Azure Container Storage Azure Arc Azure Local Microsoft Defender for Cloud Azure Monitor Microsoft Sentinel Azure Migrate View all solutions (40+) Cloud solutions for small and medium businesses Cloud migration and modernization center Data analytics for AI Azure Databases AI apps and agents Microsoft Marketplace Microsoft Sovereign Cloud AI apps and agents Responsible AI with Azure AI Infrastructure Data analytics for AI Machine learning operations (MLOps) Low-code application development on Azure Integration Services Serverless computing DevOps Migration and modernization center .NET apps migration Databases on Azure Linux on Azure Oracle on Azure SAP on the Microsoft Cloud Adaptive cloud High-performance computing (HPC) Infrastructure as a service (IaaS) Resiliency Azure Essentials Frontier Accelerate for Azure FinOps on Azure Microsoft Marketplace Azure pricing overview Create an Azure account Free Azure services Flexible purchase options Pricing calculator FinOps on Azure Maximize ROI from AI Azure savings plans Azure reservations Azure Hybrid Benefit Virtual Machines Azure SQL Microsoft Foundry Microsoft Fabric Azure Kubernetes Service (AKS) Microsoft Defender for Cloud View more Software Development Companies Microsoft Marketplace Find a partner Resources for Azure partners Get started with Azure Customer stories Analyst reports, white papers, and e-books Videos Learn more about cloud computing Documentation Explore Azure portal Developer resources Quickstart templates Resources for startups Developer community Students Azure for partners Blog Events and Webinars Learn Support Contact Sales Get started with Azure Sign in

Discover what’s coming to Microsoft Ignite, Nov 17-20, 2026. Register now.

AI infrastructure is a system, and every part of that system is connected. Decisions made in silicon and systems design influence how infrastructure is sourced, deployed, and operated across the fleet. And what we learn once that hardware is running can inform what we build next.

This feedback matters because the system never stands still. Demand shifts, component constraints emerge, new capacity comes online, and hardware requirements evolve across the fleet. Keeping infrastructure reliable means continually learning and adapting as requirements change.

At Microsoft, Azure Hardware Systems and Infrastructure works across that lifecycle, from systems architecture and design through supply chain, deployment, and fleet operations across Azure’s more than 80 regions and 500 datacenter campuses. This end-to-end view gives us an opportunity to connect insights across the hardware lifecycle, so what we learn in one part of the system can improve decisions across the others.

AI can accelerate that learning. Across our own AI transformation, we are applying agentic and AI tools to help teams connect information, understand what is changing, and act sooner while keeping human judgment at the center. The opportunity is bigger than making individual tasks faster. It’s to build a system that learns from how infrastructure is designed, sourced, and operated, and applies those learnings to what comes next. This approach is part of our broader AI transformation journey.

Start with the work, not the AI

This process has reinforced a critical lesson as we’ve scaled how we apply AI as a force multiplier across our cloud infrastructure: AI transformation starts with the work, not the technology. Speed matters, but the greater opportunity is to redesign how decisions are made: what information is available when a decision needs to happen, how quickly teams can understand what changed, and where human judgment matters most.

Our cloud supply chain is a good example of this principle in practice. Every month, demand-planning teams forecast Azure’s infrastructure needs years into the future, accounting for changing customer demand, regional needs, installed capacity, and decommissioning activity. When a plan changes, determining why could require reconciling information across multiple systems, turning a single investigation into a lengthy process. It was tempting to look at that work and ask where we could add an agent.

For any company, putting AI on top of a fragmented process can simply make the fragmentation move faster. Before applying AI, our teams mapped and simplified the work, established a shared data foundation with quality, governance, and access controls, and identified decisions where people needed to remain accountable. We call this approach “Lean before AI.”

Starting with end-to-end processes and taking an AI-driven approach provides new ways of working and moves teams to parallel execution rather than sequential handoffs, resulting in integrated, collaborative workflows.

From days of research to decisions in minutes

With this foundation in place, we approached demand planning differently. A multi-agent workflow can examine signals such as installed-base shifts, regional demand, and decommissioning changes, then help planners understand what changed, where, and what drove the movement. Work that previously took five to seven business days can now be completed in hours, and sometimes in less than 20 minutes. Across more than five monthly planning cycles, our full demand-planning team saw approximately 50% less manual effort and cycle time fell by up to 75% in selected workflows.

The same pattern is taking shape across planning, product data, sourcing, fulfillment, logistics, and operational workflows. Specialized agents are helping teams spend less time finding and reconciling information and more time applying expertise.

In fulfillment, understanding why rack delivery is blocked from meeting customer demand could require teams to pull information manually from multiple sources. An intelligent assistant now brings together information about blockers and compatible or incompatible supplies, saving investigation time by as much as 55%. In logistics, an AI-powered logistics agent brings together data across air, land, and sea options to help teams evaluate speed, cost, as well as carbon tradeoffs and forecast emissions.

Building the learning loop

These individual applications matter, but the larger opportunity is to connect them. Our cloud supply chain team is moving toward end-to-end multi-agent workflows across bill-of-materials generation, capacity delivery, spare-parts management, capacity docking, and sales and operations execution. This work reflects a broader shift in our business: moving beyond isolated experiments toward a faster, more resilient, and intelligent operating system that places human judgment at the center.

The important outcome is not simply speed. Planners can begin with connected evidence instead of spending days assembling it, giving them more time to examine the explanation, add business context, and determine what it means for the decision ahead.

That is the learning loop we want. AI helps people reach the evidence faster. People bring context and judgment, act on what they learn, and create new information that can improve the next decision.

Learning across the fleet

The hardware lifecycle does not end when a server reaches a datacenter. Once infrastructure is deployed, the challenge becomes keeping it operating reliably for customers. Across millions of nodes in our fleet, continuous monitoring generates signals that help our teams investigate issues and determine root causes to decide how to act.

Across Azure, we’re moving cloud reliability upstream—transforming fleet management from reactive firefighting into a closed-loop system that prevents defects, predicts failures, and automatically restores hardware back into service. As we move toward a self-healing fleet, we are applying the same principles: connecting data across the lifecycle, continuous evaluation, redesigning workflows for human-agent orchestration, and keeping engineers in control of production decisions. Ultimately, this is also when the next learning cycle begins: systems collect and analyze information about the fleet, and failure patterns become evidence for how suppliers design and build the next generation of hardware.

Azure failure prediction and detection uses AI to analyze fleet telemetry so teams can identify emerging hardware failure patterns sooner and take action before potential issues impact customers. Engineers retain oversight of production decisions. This has already resulted in a 92% reduction in disk-related virtual machine (VM) interruptions and reduced repair time on out-of-service nodes by 53%. For rack managers, prediction provides up to three days of advance warning, enabling proactive recovery that cuts out-of-service repairs by 40%.

We also proactively and periodically screen our fleet to identify hardware vulnerable to silent data corruption before customer workloads are deployed, helping strengthen platform reliability.

We are also developing workflows that preserve context as hardware moves through investigation and recovery. For faulted resources, these workflows track assignment, action, outcome, and next step, with policy and approval controls around fleet actions. History and outcomes can then inform future decisions.

This is where the broader systems story comes together. Demand-planning decisions influence sourcing. Logistics affects when capacity reaches a datacenter. Once hardware is running, fleet telemetry and operational outcomes create another source of learning. Information about component performance can help teams proactively address potential issues before customers experience them. Those insights can also feed forward into the next generation of silicon, system, and rack design, while giving suppliers information to improve future components.

AI can make that loop faster, but the value comes from helping the system work better as a whole.

What we learned when things did not work

Some of our most useful lessons came from approaches that fell short.

We learned that applying AI to one part of a fragmented process can accelerate that task while creating more work somewhere else. An agent might produce its output faster, but if a downstream team must interpret, reformat or reconcile it manually, the workflow as a whole has not improved.

That changed how we measured success. Instead of evaluating only the task an agent performs, teams must examine the full workflow: the work removed, the new work created, the quality of the decision, and the outcome.

Reliable, accessible, and well-governed data is a prerequisite, not an afterthought. AI cannot compensate for conflicting definitions, unclear permissions, or information isolated across systems.

And we learned not to become attached to a particular architecture or agent. Models, frameworks, and business needs continue to change. A solution that is useful today may need to be redesigned six months from now or retired if the original business need no longer applies.

The resulting rhythm is practical: begin with a consequential decision, simplify the work around it, connect the right governed data, build alongside the people who know the work, evaluate the complete outcome, and keep changing as the business and technology evolve.

What we measure next

The next phase of enterprise AI will require us to measure more than adoption: how many people use an agent, how many agents are deployed, or how much time they save. Those measures matter, but they don’t tell us whether the work itself has improved. Can a planner understand in minutes a change that once took days to explain? Can a fulfillment manager resolve a capacity blocker without manually reconciling information across systems? Can an engineer identify a potential hardware problem before it becomes a customer problem? And can teams apply their expertise, so the next decision is better than the last?

That is the direction we are pursuing: a way of working that keeps learning as the technology and the business change. We are not trying to build the largest collection of agents or automate decisions simply because we can. We are working to produce more useful output from the infrastructure, information, and expertise already in the system.

If AI is going to change what the world can build, we need to keep evolving how we build the infrastructure behind it. By connecting insights across silicon, systems, supply chain, and fleet operations, each stage can help improve the next. The result is reliable infrastructure ready when customers need it, and a system that learns how to deliver it better with every cycle.

Azure infrastructure

Learn more about Microsoft’s systems approach.

Two people having a conversation while walking outside.

WE ARE MICROSOFT

Explore Microsoft Foundry

The future of AI starts here. Envision your next great AI app with the latest technologies. Get started with Azure.