Azure Cost Management and Service Level Agreements for AZ-900 Candidates
Introduction: Balancing Cost and Availability in Azure
For AZ-900, you need to understand more than service names. That’s the real game here — understanding trade-offs. In Azure, cost, performance, and availability are linked: a cheaper design may reduce resilience, while a more available design usually adds cost through redundancy, higher tiers, or more complex architecture.
Azure Cost Management helps you estimate, monitor, and control spend. Service Level Agreements, or SLAs, are Microsoft’s promise about availability for a specific service, as long as the service is deployed and used under the stated conditions. The exam often tests whether you can choose the right tool or design based on a business requirement: lower cost, better visibility, higher uptime, or migration planning.
AZ-900 must-know: Pricing Calculator = estimate future Azure cost. The TCO Calculator is for comparing your on-premises costs with Azure. Cost Management = analyze actual spend. SLA means Microsoft’s availability promise for a service — it doesn’t mean your whole application is guaranteed to stay up from end to end.
How Azure Pricing Works
Azure pricing is consumption-based, which basically means you pay for what you use instead of shelling out for hardware before you even switch anything on. That’s a pretty big shift from capital expenditure, or CapEx, to operational expenditure, or OpEx. Charges are usually based on meters: compute runtime, storage capacity, transactions, bandwidth, and service-specific features.
Pricing varies by region, service tier, SKU, operating system, licensing model, redundancy option, and sometimes transaction class. Taxes, currency, exchange rates, support plans, marketplace items, and billing agreement type can also affect the final invoice.
For virtual machines, Azure VM compute is generally billed per second for many VM resources after the first minute while the VM is running, but billing behavior varies by resource type and offer. Always verify the pricing details for the specific service and configuration.
A common exam and real-world trap is VM shutdown behavior. If you want to stop VM compute charges, the VM needs to be deallocated. If you just shut down the guest operating system inside the VM, it can still sit there in a stopped-but-allocated state, and yep, that can keep the compute meter running. Even after deallocation, other attached resources may still cost money.
| Resource state/item | Compute billed? | Other charges may continue? |
|---|---|---|
| VM running | Yes | Yes - disks, backups, IPs, networking |
| VM stopped in guest OS only | May still be yes | Yes |
| VM deallocated | No compute charge | Yes - OS disk, data disks, snapshots, backups, some IPs |
| VM deleted | No | Only if dependent resources were kept |
Important: If you want to reduce VM cost, think deallocate or delete, not just “stop.”
Azure also offers different purchasing models:
- Pay-as-you-go: no long-term commitment.
- Reservations: discounts for specific eligible services/resources, commonly for 1-year or 3-year terms.
- Azure savings plan for compute: discounted rates in exchange for a fixed hourly spend commitment for 1 or 3 years on eligible compute services.
- Azure Hybrid Benefit: use eligible Windows Server or SQL Server licenses to reduce cloud cost.
Reservations are more targeted and are often best when you know the resource family and usage pattern. Savings plans are more flexible across eligible compute usage, but they do not apply to all Azure services such as every storage or bandwidth charge.
Network pricing also matters. Inbound data transfer is generally free, while outbound internet egress is commonly charged. Inter-region data transfer can also add to the bill, depending on where the data starts, where it ends up, and which Azure service is shifting it around. Storage pricing isn’t just about how much space you’ve got. It can also include transactions, redundancy, access tier, and even how often you pull the data back out.
| Cost factor | Example | Why it changes the bill |
|---|---|---|
| Compute runtime | VM runs 24/7 | More runtime means more compute cost |
| Storage capacity | 500 GB managed disk | More stored data increases monthly cost |
| Storage transactions | Frequent reads/writes | Some storage services charge per operation |
| Network egress | Data sent to internet | Outbound traffic is commonly billed |
| SKU/tier | Premium disk vs standard | Higher performance/features cost more |
| Region | Different Azure locations | Regional pricing varies |
| Licensing | Windows Server, SQL Server | Software licensing changes total cost |
| Billing offer | Pay-As-You-Go, MCA, EA, CSP | Discounts and billing structure can differ |
| Support plan | Higher support tier | Billed separately from resource consumption |
Cost Estimation, Comparison, and Monitoring Tools
The exam loves checking whether you can match the right tool to the right problem.
| Tool | Purpose | When to use | Example output |
|---|---|---|---|
| Azure Pricing Calculator | Estimate planned Azure cost | Before deployment | Estimated monthly cost for VMs, storage, bandwidth |
| Azure TCO Calculator | Compare on-premises with Azure | Migration planning | 3-year on-prem vs Azure cost comparison |
| Azure Cost Management + Billing | Analyze actual spend and billing | After deployment | Cost analysis, budgets, alerts, exports, invoices |
Pricing Calculator is for planned Azure solutions. You plug in things like region, VM size, runtime assumptions, storage type, and bandwidth estimates. It is an estimate, not your live bill.
TCO Calculator is for migration business cases. The TCO Calculator compares datacenter costs — things like hardware, power, cooling, facilities, and admin effort — against Azure estimates. It’s useful, absolutely, but the output is only as good as the assumptions you feed into it, so it’s not a guarantee.
Cost Management + Billing is the portal experience used to analyze current spend, view billing data, create budgets, forecast trends, and export usage. Cost Management is the cost analysis and optimization part of that broader billing experience.
A practical workflow in Cost Management looks like this:
- Choose the right scope: management group, subscription, or resource group.
- Open Cost analysis and filter by service name, resource group, location, meter, or tag.
- Compare current spend with the previous month.
- Create a budget with thresholds such as 50%, 80%, and 100%.
- Configure alert notifications so owners know early.
- Export cost and usage data to storage for finance or BI reporting if needed.
Budgets and alerts are useful, but they do not automatically shut down resources. They’re visibility and notification tools, not automatic spend blockers.
Azure Advisor can sit alongside Cost Management and point out things like rightsizing opportunities, idle resources, or other cost-saving ideas. I like to think of Advisor as the recommendation engine and Cost Management as the place where you actually see, track, and govern the spend.
Cost Optimization and Governance
Optimization means aligning spend with business value, not just cutting cost. Common methods include rightsizing, deallocating unused VMs, deleting orphaned resources, autoscaling, reservations, savings plans, Azure Hybrid Benefit, and Azure Spot Virtual Machines for workloads that can tolerate interruption.
| Method | Best for | Trade-offs |
|---|---|---|
| Rightsizing | Oversized workloads | Too much downsizing can hurt performance |
| Deallocate/delete unused resources | Dev/test and temporary assets | Must confirm what can be safely removed |
| Autoscaling | Variable demand | More design complexity |
| Reservations | Predictable eligible resources | Less flexible than pay-as-you-go |
| Azure savings plan for compute | Predictable eligible compute usage | Requires hourly spend commitment |
| Azure Hybrid Benefit | Eligible Windows/SQL workloads | Requires license compliance checks |
| Azure Spot VMs | Interruptible batch/test workloads | Eviction can occur due to capacity or price constraints |
| Budgets/alerts | Spend visibility | Do not directly stop billing |
| Tags | Chargeback/showback | Only valuable if consistently applied |
Use reservations when a specific eligible workload is stable and long-lived. Use Azure savings plan for compute when you want discount flexibility across eligible compute services. Use Spot VMs only when interruption is acceptable.
Governance is what makes cost visible and, just as importantly, controllable. Azure organises resources through management groups, subscriptions, resource groups, and then the actual resources themselves. Cost views and permissions depend on scope and role assignment, so least-privilege access really matters. Teams often need read access to cost data without having broad rights to create or change resources.
Tags are central to chargeback and showback, but they aren’t automatically inherited by every resource in every situation. That means governance automation matters. A practical tag schema might include CostCenter, App, Env, and Owner.
Azure Policy helps cut down on accidental overspend by enforcing deployment standards. Common examples include:
- Require specific tags
- Restrict allowed regions
- Restrict allowed VM SKUs
- Audit or deny public IP creation in dev environments
Policy effects can include deny, audit, append, or modify depending on the definition. Policy doesn’t enforce budgets or stop charges directly, but it can absolutely help prevent expensive or noncompliant deployments in the first place.
Troubleshooting cost spikes checklist: confirm the scope, filter by service and resource group, check tags, review region or SKU changes, verify whether VMs were deallocated, look for orphaned disks/snapshots/public IPs, inspect outbound data transfer, and review backup or log growth. Cost anomaly alerts can help you spot unusual patterns early, but they don’t replace proper investigation.
Storage, Network, and Performance Trade-offs
For AZ-900, it’s worth remembering that performance choices often affect both cost and availability. A higher SKU might give you more CPU, memory, IOPS, throughput, or lower latency — but, naturally, it’ll usually cost more too. Undersizing can look cheap on paper, but if the app performs badly or keeps needing emergency scaling, those ‘savings’ disappear fast.
Storage costs depend on capacity, access tier, redundancy, and transactions. Hot, cool, and archive tiers are basically a trade-off between speed, retrieval behavior, and price. Redundancy options like LRS, ZRS, GRS, and GZRS affect both resilience and cost, so they’re not just technical choices — they’re budget choices too. More copies generally mean higher cost but better durability or regional resilience.
Network cost surprises often come from data movement. Internet egress, inter-region transfer, VPN Gateway, ExpressRoute, and even some load-balancing-related traffic can all contribute to the bill. The exam usually keeps this conceptual, but the principle is important: moving data isn’t always free.
Understanding Azure SLAs and Downtime
An SLA is Microsoft’s availability commitment for a specific service over a defined measurement period, usually monthly. It’s not the same as guaranteed application uptime, backup, disaster recovery, or business continuity. Some services in preview, free tiers, or certain configurations may not have an SLA at all, or the SLA may only apply if you deploy the service in a particular supported way.
If Microsoft misses an SLA, the usual remedy is service credits, subject to SLA terms. Service credits aren’t compensation for your business loss, though — that’s an important distinction.
Approximate downtime is a lot easier to understand when you convert the percentages into actual time. The table below assumes a 30-day month.
| SLA % | Approx monthly downtime | Approx yearly downtime |
|---|---|---|
| 99% | About 7 hours 18 minutes | About 3 days 15 hours 36 minutes |
| 99.9% | About 43 minutes 49 seconds | About 8 hours 45 minutes 36 seconds |
| 99.95% | About 21 minutes 54 seconds | About 4 hours 22 minutes 48 seconds |
| 99.99% | About 4 minutes 23 seconds | About 52 minutes 35 seconds |
Even tiny percentage differences can matter a lot in the real world. “One more nine” can mean a major reduction in maximum downtime implied by the availability target.
Also distinguish SLA from SLO. The SLA is the external commitment; the SLO, or Service Level Objective, is the internal target. For AZ-900, focus mainly on what the SLA means and what it means for the business.
Composite SLA and Availability Design
Composite SLA comes into play when your application depends on multiple services. For serial dependencies, a simple way to estimate overall availability is to multiply the component SLAs together. For example:
99.9% × 99.9% = 99.8001%
That means two required dependent services can produce lower overall availability than either service alone. The same idea extends to more components: more serial dependencies often reduce effective availability.
But that multiplication logic applies to dependent serial components assuming independence and no redundancy. Redundant parallel components change the outcome. On the other hand, two application instances behind a load balancer can improve effective availability compared with a single instance, because one instance can keep serving traffic if the other one fails.
For Azure VM-based designs, availability options include:
| Design | Use case | Cost/complexity | Resilience note |
|---|---|---|---|
| Single instance | Low-criticality dev/test | Lowest | Single point of failure |
| Availability Set | VMs in one datacenter context | Higher | Distributes VMs across fault and update domains |
| Availability Zones | Stronger in-region resilience | Higher | Uses physically separate datacenter locations in a region |
| Multi-region DR | Regional outage planning | Highest | Requires replication and failover design |
Two VMs on their own don’t automatically create application availability. You’ll usually need some kind of traffic distribution or failover component, like Azure Load Balancer or Application Gateway. Region pairs are useful for disaster recovery planning, but they don’t automatically fail over your application. You’ve got to design the replication, failover, and recovery process properly.
High availability, fault tolerance, and disaster recovery are related, but they’re not the same thing. High availability is about staying online through common failures. Fault tolerance goes a step further and tries to keep the system running even when a component fails. Disaster recovery is about restoring service after bigger failures, and you’ll often hear it discussed in terms of RTO and RPO.
AZ-900 Scenario Review
| Business need | Best Azure choice | Cost effect | Exam lesson |
|---|---|---|---|
| Estimate a new Azure deployment | Pricing Calculator | No live billing impact | Estimate future Azure cost |
| Compare datacenter costs with Azure | TCO Calculator | Supports migration planning | Compare on-prem vs cloud |
| Analyze rising spend in a live subscription | Cost Management + Billing | Improves visibility and control | Use actual spend tools after deployment |
| Stable 24/7 eligible compute workload | Reservation or savings plan | Can reduce cost | Commitment discounts do not improve SLA |
| Allocate costs by department | Tags plus Cost Management | Improves accountability | Tags help cost allocation, not availability |
| Improve web app uptime | Multiple instances plus Load Balancer, possibly zones | Higher | More resilience usually means more cost |
| Use a preview service for production | Check service terms carefully | Varies | Preview may have no SLA |
Common AZ-900 Pitfalls and Final Review
The most common mistakes are predictable:
- Confusing Pricing Calculator, TCO Calculator, and Cost Management
- Assuming SLA equals guaranteed application uptime
- Forgetting that a VM must be deallocated to stop compute billing
- Assuming tags reduce cost directly
- Thinking budgets automatically stop resources
- Believing reservations or savings plans improve availability
- Ignoring that some services/configurations may not have an SLA
Exam memory aid: Plan = Pricing Calculator. Compare = TCO Calculator. Analyze = Cost Management. Optimize = rightsize, deallocate, reserve, govern. Improve uptime = redundancy, load balancing, zones, and DR design.
Do not try to memorize every price. For AZ-900, focus on purpose, trade-offs, and correct tool selection. If a question mentions billing visibility, think budgets, tags, scopes, and Cost Management. If it mentions uptime, think SLA, architecture, and redundancy. If it mentions both, the answer is usually about balancing business requirements rather than choosing the cheapest option.
Final takeaway: Azure success is not about finding the lowest price or the highest SLA in isolation. It is about matching cost, performance, and availability to the workload’s real business need.