Azure Cost Management and Service Level Agreements for AZ-900 Candidates

Introduction: Balancing Cost and Availability in Azure

For AZ-900, you need to understand more than service names. That’s the real game here — understanding trade-offs. In Azure, cost, performance, and availability are linked: a cheaper design may reduce resilience, while a more available design usually adds cost through redundancy, higher tiers, or more complex architecture.

Azure Cost Management helps you estimate, monitor, and control spend. Service Level Agreements, or SLAs, are Microsoft’s promise about availability for a specific service, as long as the service is deployed and used under the stated conditions. The exam often tests whether you can choose the right tool or design based on a business requirement: lower cost, better visibility, higher uptime, or migration planning.

AZ-900 must-know: Pricing Calculator = estimate future Azure cost. The TCO Calculator is for comparing your on-premises costs with Azure. Cost Management = analyze actual spend. SLA means Microsoft’s availability promise for a service — it doesn’t mean your whole application is guaranteed to stay up from end to end.

How Azure Pricing Works

Azure pricing is consumption-based, which basically means you pay for what you use instead of shelling out for hardware before you even switch anything on. That’s a pretty big shift from capital expenditure, or CapEx, to operational expenditure, or OpEx. Charges are usually based on meters: compute runtime, storage capacity, transactions, bandwidth, and service-specific features.

Pricing varies by region, service tier, SKU, operating system, licensing model, redundancy option, and sometimes transaction class. Taxes, currency, exchange rates, support plans, marketplace items, and billing agreement type can also affect the final invoice.

For virtual machines, Azure VM compute is generally billed per second for many VM resources after the first minute while the VM is running, but billing behavior varies by resource type and offer. Always verify the pricing details for the specific service and configuration.

A common exam and real-world trap is VM shutdown behavior. If you want to stop VM compute charges, the VM needs to be deallocated. If you just shut down the guest operating system inside the VM, it can still sit there in a stopped-but-allocated state, and yep, that can keep the compute meter running. Even after deallocation, other attached resources may still cost money.

Resource state/itemCompute billed?Other charges may continue?
VM runningYesYes - disks, backups, IPs, networking
VM stopped in guest OS onlyMay still be yesYes
VM deallocatedNo compute chargeYes - OS disk, data disks, snapshots, backups, some IPs
VM deletedNoOnly if dependent resources were kept

Important: If you want to reduce VM cost, think deallocate or delete, not just “stop.”

Azure also offers different purchasing models:

  • Pay-as-you-go: no long-term commitment.
  • Reservations: discounts for specific eligible services/resources, commonly for 1-year or 3-year terms.
  • Azure savings plan for compute: discounted rates in exchange for a fixed hourly spend commitment for 1 or 3 years on eligible compute services.
  • Azure Hybrid Benefit: use eligible Windows Server or SQL Server licenses to reduce cloud cost.

Reservations are more targeted and are often best when you know the resource family and usage pattern. Savings plans are more flexible across eligible compute usage, but they do not apply to all Azure services such as every storage or bandwidth charge.

Network pricing also matters. Inbound data transfer is generally free, while outbound internet egress is commonly charged. Inter-region data transfer can also add to the bill, depending on where the data starts, where it ends up, and which Azure service is shifting it around. Storage pricing isn’t just about how much space you’ve got. It can also include transactions, redundancy, access tier, and even how often you pull the data back out.

Cost factorExampleWhy it changes the bill
Compute runtimeVM runs 24/7More runtime means more compute cost
Storage capacity500 GB managed diskMore stored data increases monthly cost
Storage transactionsFrequent reads/writesSome storage services charge per operation
Network egressData sent to internetOutbound traffic is commonly billed
SKU/tierPremium disk vs standardHigher performance/features cost more
RegionDifferent Azure locationsRegional pricing varies
LicensingWindows Server, SQL ServerSoftware licensing changes total cost
Billing offerPay-As-You-Go, MCA, EA, CSPDiscounts and billing structure can differ
Support planHigher support tierBilled separately from resource consumption

Cost Estimation, Comparison, and Monitoring Tools

The exam loves checking whether you can match the right tool to the right problem.

ToolPurposeWhen to useExample output
Azure Pricing CalculatorEstimate planned Azure costBefore deploymentEstimated monthly cost for VMs, storage, bandwidth
Azure TCO CalculatorCompare on-premises with AzureMigration planning3-year on-prem vs Azure cost comparison
Azure Cost Management + BillingAnalyze actual spend and billingAfter deploymentCost analysis, budgets, alerts, exports, invoices

Pricing Calculator is for planned Azure solutions. You plug in things like region, VM size, runtime assumptions, storage type, and bandwidth estimates. It is an estimate, not your live bill.

TCO Calculator is for migration business cases. The TCO Calculator compares datacenter costs — things like hardware, power, cooling, facilities, and admin effort — against Azure estimates. It’s useful, absolutely, but the output is only as good as the assumptions you feed into it, so it’s not a guarantee.

Cost Management + Billing is the portal experience used to analyze current spend, view billing data, create budgets, forecast trends, and export usage. Cost Management is the cost analysis and optimization part of that broader billing experience.

A practical workflow in Cost Management looks like this:

  1. Choose the right scope: management group, subscription, or resource group.
  2. Open Cost analysis and filter by service name, resource group, location, meter, or tag.
  3. Compare current spend with the previous month.
  4. Create a budget with thresholds such as 50%, 80%, and 100%.
  5. Configure alert notifications so owners know early.
  6. Export cost and usage data to storage for finance or BI reporting if needed.

Budgets and alerts are useful, but they do not automatically shut down resources. They’re visibility and notification tools, not automatic spend blockers.

Azure Advisor can sit alongside Cost Management and point out things like rightsizing opportunities, idle resources, or other cost-saving ideas. I like to think of Advisor as the recommendation engine and Cost Management as the place where you actually see, track, and govern the spend.

Cost Optimization and Governance

Optimization means aligning spend with business value, not just cutting cost. Common methods include rightsizing, deallocating unused VMs, deleting orphaned resources, autoscaling, reservations, savings plans, Azure Hybrid Benefit, and Azure Spot Virtual Machines for workloads that can tolerate interruption.

MethodBest forTrade-offs
RightsizingOversized workloadsToo much downsizing can hurt performance
Deallocate/delete unused resourcesDev/test and temporary assetsMust confirm what can be safely removed
AutoscalingVariable demandMore design complexity
ReservationsPredictable eligible resourcesLess flexible than pay-as-you-go
Azure savings plan for computePredictable eligible compute usageRequires hourly spend commitment
Azure Hybrid BenefitEligible Windows/SQL workloadsRequires license compliance checks
Azure Spot VMsInterruptible batch/test workloadsEviction can occur due to capacity or price constraints
Budgets/alertsSpend visibilityDo not directly stop billing
TagsChargeback/showbackOnly valuable if consistently applied

Use reservations when a specific eligible workload is stable and long-lived. Use Azure savings plan for compute when you want discount flexibility across eligible compute services. Use Spot VMs only when interruption is acceptable.

Governance is what makes cost visible and, just as importantly, controllable. Azure organises resources through management groups, subscriptions, resource groups, and then the actual resources themselves. Cost views and permissions depend on scope and role assignment, so least-privilege access really matters. Teams often need read access to cost data without having broad rights to create or change resources.

Tags are central to chargeback and showback, but they aren’t automatically inherited by every resource in every situation. That means governance automation matters. A practical tag schema might include CostCenter, App, Env, and Owner.

Azure Policy helps cut down on accidental overspend by enforcing deployment standards. Common examples include:

  • Require specific tags
  • Restrict allowed regions
  • Restrict allowed VM SKUs
  • Audit or deny public IP creation in dev environments

Policy effects can include deny, audit, append, or modify depending on the definition. Policy doesn’t enforce budgets or stop charges directly, but it can absolutely help prevent expensive or noncompliant deployments in the first place.

Troubleshooting cost spikes checklist: confirm the scope, filter by service and resource group, check tags, review region or SKU changes, verify whether VMs were deallocated, look for orphaned disks/snapshots/public IPs, inspect outbound data transfer, and review backup or log growth. Cost anomaly alerts can help you spot unusual patterns early, but they don’t replace proper investigation.

Storage, Network, and Performance Trade-offs

For AZ-900, it’s worth remembering that performance choices often affect both cost and availability. A higher SKU might give you more CPU, memory, IOPS, throughput, or lower latency — but, naturally, it’ll usually cost more too. Undersizing can look cheap on paper, but if the app performs badly or keeps needing emergency scaling, those ‘savings’ disappear fast.

Storage costs depend on capacity, access tier, redundancy, and transactions. Hot, cool, and archive tiers are basically a trade-off between speed, retrieval behavior, and price. Redundancy options like LRS, ZRS, GRS, and GZRS affect both resilience and cost, so they’re not just technical choices — they’re budget choices too. More copies generally mean higher cost but better durability or regional resilience.

Network cost surprises often come from data movement. Internet egress, inter-region transfer, VPN Gateway, ExpressRoute, and even some load-balancing-related traffic can all contribute to the bill. The exam usually keeps this conceptual, but the principle is important: moving data isn’t always free.

Understanding Azure SLAs and Downtime

An SLA is Microsoft’s availability commitment for a specific service over a defined measurement period, usually monthly. It’s not the same as guaranteed application uptime, backup, disaster recovery, or business continuity. Some services in preview, free tiers, or certain configurations may not have an SLA at all, or the SLA may only apply if you deploy the service in a particular supported way.

If Microsoft misses an SLA, the usual remedy is service credits, subject to SLA terms. Service credits aren’t compensation for your business loss, though — that’s an important distinction.

Approximate downtime is a lot easier to understand when you convert the percentages into actual time. The table below assumes a 30-day month.

SLA %Approx monthly downtimeApprox yearly downtime
99%About 7 hours 18 minutesAbout 3 days 15 hours 36 minutes
99.9%About 43 minutes 49 secondsAbout 8 hours 45 minutes 36 seconds
99.95%About 21 minutes 54 secondsAbout 4 hours 22 minutes 48 seconds
99.99%About 4 minutes 23 secondsAbout 52 minutes 35 seconds

Even tiny percentage differences can matter a lot in the real world. “One more nine” can mean a major reduction in maximum downtime implied by the availability target.

Also distinguish SLA from SLO. The SLA is the external commitment; the SLO, or Service Level Objective, is the internal target. For AZ-900, focus mainly on what the SLA means and what it means for the business.

Composite SLA and Availability Design

Composite SLA comes into play when your application depends on multiple services. For serial dependencies, a simple way to estimate overall availability is to multiply the component SLAs together. For example:

99.9% × 99.9% = 99.8001%

That means two required dependent services can produce lower overall availability than either service alone. The same idea extends to more components: more serial dependencies often reduce effective availability.

But that multiplication logic applies to dependent serial components assuming independence and no redundancy. Redundant parallel components change the outcome. On the other hand, two application instances behind a load balancer can improve effective availability compared with a single instance, because one instance can keep serving traffic if the other one fails.

For Azure VM-based designs, availability options include:

DesignUse caseCost/complexityResilience note
Single instanceLow-criticality dev/testLowestSingle point of failure
Availability SetVMs in one datacenter contextHigherDistributes VMs across fault and update domains
Availability ZonesStronger in-region resilienceHigherUses physically separate datacenter locations in a region
Multi-region DRRegional outage planningHighestRequires replication and failover design

Two VMs on their own don’t automatically create application availability. You’ll usually need some kind of traffic distribution or failover component, like Azure Load Balancer or Application Gateway. Region pairs are useful for disaster recovery planning, but they don’t automatically fail over your application. You’ve got to design the replication, failover, and recovery process properly.

High availability, fault tolerance, and disaster recovery are related, but they’re not the same thing. High availability is about staying online through common failures. Fault tolerance goes a step further and tries to keep the system running even when a component fails. Disaster recovery is about restoring service after bigger failures, and you’ll often hear it discussed in terms of RTO and RPO.

AZ-900 Scenario Review

Business needBest Azure choiceCost effectExam lesson
Estimate a new Azure deploymentPricing CalculatorNo live billing impactEstimate future Azure cost
Compare datacenter costs with AzureTCO CalculatorSupports migration planningCompare on-prem vs cloud
Analyze rising spend in a live subscriptionCost Management + BillingImproves visibility and controlUse actual spend tools after deployment
Stable 24/7 eligible compute workloadReservation or savings planCan reduce costCommitment discounts do not improve SLA
Allocate costs by departmentTags plus Cost ManagementImproves accountabilityTags help cost allocation, not availability
Improve web app uptimeMultiple instances plus Load Balancer, possibly zonesHigherMore resilience usually means more cost
Use a preview service for productionCheck service terms carefullyVariesPreview may have no SLA

Common AZ-900 Pitfalls and Final Review

The most common mistakes are predictable:

  • Confusing Pricing Calculator, TCO Calculator, and Cost Management
  • Assuming SLA equals guaranteed application uptime
  • Forgetting that a VM must be deallocated to stop compute billing
  • Assuming tags reduce cost directly
  • Thinking budgets automatically stop resources
  • Believing reservations or savings plans improve availability
  • Ignoring that some services/configurations may not have an SLA

Exam memory aid: Plan = Pricing Calculator. Compare = TCO Calculator. Analyze = Cost Management. Optimize = rightsize, deallocate, reserve, govern. Improve uptime = redundancy, load balancing, zones, and DR design.

Do not try to memorize every price. For AZ-900, focus on purpose, trade-offs, and correct tool selection. If a question mentions billing visibility, think budgets, tags, scopes, and Cost Management. If it mentions uptime, think SLA, architecture, and redundancy. If it mentions both, the answer is usually about balancing business requirements rather than choosing the cheapest option.

Final takeaway: Azure success is not about finding the lowest price or the highest SLA in isolation. It is about matching cost, performance, and availability to the workload’s real business need.