Back to home

Cloud

My perspective on governing cloud platforms across different environments.

Cloud terms used on this page
CAF
Cloud Adoption Framework: guidance for an organization's cloud journey.
WAF
Well-Architected Framework here—not a web application firewall. Guidance for workload design and operation.
Landing zone
A shared cloud foundation for accounts, identity, networking, and governance.
SLO / RTO / RPO
Service-level objective; recovery-time objective; recovery-point objective. Targets for service quality, recovery time, and acceptable data loss.

How I approach cloud governance

Beyond moving workloads to the cloud: governance that continues through operations.

I have developed a mature governance practice across GCP, AWS, Azure, Alibaba Cloud, and OpenStack. Across platforms, I focus on an approach that can be put into practice and refined: define business goals and ownership, establish platform baselines, then improve through policy, security, cost, and operational feedback.

Cloud Adoption Framework (Microsoft Learn, opens in a new tab)

I draw on its path from Strategy, Plan, and Ready through Govern, Secure, and Manage to connect business goals, platform readiness, and ongoing governance. The methodology informs multi-cloud practice without copying Azure-specific implementation details.

Azure Well-Architected Framework (Microsoft Learn, opens in a new tab)

I review workloads through its five pillars—Reliability, Security, Cost Optimization, Operational Excellence, and Performance Efficiency—balancing risk, experience, and investment.

These frameworks help frame the questions; implementation depends on each platform’s capabilities, organizational constraints, and workload needs.

How organizations adopt cloud

Based on the Cloud Adoption Framework (Microsoft Learn, opens in a new tab). It is Azure guidance; the organizational questions also inform adoption on AWS, GCP, Alibaba Cloud, and private cloud.

CAF connects business goals, organizational preparation, and cloud delivery. Strategy, Plan, Ready, and Adopt provide an initial path; Govern, Secure, and Manage are ongoing practices planned early and applied throughout adoption and operations. Revisit earlier decisions as needs change. The examples below follow a hypothetical customer portal.

Microsoft Cloud Adoption Framework: Strategy, Plan, Ready, and Adopt in sequence; Govern, Secure, and Manage as ongoing processes. Azure Architecture Center and the Well-Architected Framework support implementation and workload design. View full-size diagram (opens in a new tab)
Microsoft’s Cloud Adoption Framework overview. Source: Microsoft Learn (opens in a new tab). View full-size diagram (opens in a new tab).
Read the diagram in text

Cloud adoption · sequential

  1. 01Strategy
  2. 02Plan
  3. 03Ready
  4. 04Adopt

Operations · in parallel

  • Govern
  • Secure
  • Manage

Azure Architecture Center provides implementation guidance; the Well-Architected Framework provides principles for workload design. Both support readiness, adoption, and ongoing operations.

  1. Strategy

    Connect cloud investment to business outcomes before choosing services. Bring business, finance, IT, and security leaders together to agree on motivations, scope, risk tolerance, and measurable success criteria.

    Outcome
    A business case with an accountable sponsor, investment priorities, and success measures.
    Example
    For a customer portal, prioritize faster releases and seasonal capacity. Measure deployment lead time and successful customer requests against the current baseline.
  2. Plan

    Turn the strategy into an achievable roadmap. Inventory applications, data, and dependencies; assess cloud readiness and skills; estimate costs; and assign responsibilities between platform and workload teams.

    Outcome
    A prioritized backlog, migration waves, budget, training plan, and clear ownership.
    Example
    Map the portal’s database, identity provider, and payment dependencies. Pilot a low-risk environment first, then schedule production migration around business constraints.
  3. Ready

    Build a repeatable landing zone before teams deploy workloads. Separate shared platform services from workload environments, and establish identity, network connectivity, account structure, policy, logging, and billing baselines through infrastructure as code.

    Outcome
    A tested cloud foundation that workload teams can use within agreed guardrails.
    Example
    Provide separate production and nonproduction subscriptions, private database connectivity, centralized logs, and owner and cost-center tags through a reusable Terraform template.
  4. Adopt

    Migrate, modernize, or build workloads according to business value and technical constraints. Choose the approach per application, test functionality and performance, and plan data transfer, cutover, rollback, and operational handover. Repeat as needs evolve.

    Outcome
    A production workload that meets acceptance criteria and has a named operating team.
    Example
    Move the portal with a rehearsed database cutover and rollback plan. Modernize background jobs to a managed queue when that change delivers enough value to justify the effort.
  5. Govern

    Translate business risks and compliance obligations into enforceable policies. Define allowed regions, resource ownership, spending controls, and exception handling; review compliance continuously as the estate grows.

    Outcome
    Documented policies, compliance evidence, and a process for time-limited exceptions.
    Example
    Require owner tags, restrict deployment to approved regions, and deny public access to storage. Route budget alerts to workload owners; alerts alone do not cap spending.
  6. Secure

    Protect the environment across its lifecycle using Zero Trust: verify explicitly, grant least privilege, and assume breach. Combine identity controls, network segmentation, data protection, vulnerability management, threat detection, and incident response.

    Outcome
    A security baseline, prioritized remediation work, and tested response procedures.
    Example
    Require multifactor authentication for administrators, use managed workload identities instead of embedded keys, keep the database private, and investigate suspicious sign-ins through centralized security monitoring.
  7. Manage

    Keep cloud services healthy after deployment. Agree on service and recovery targets, monitor user experience and platform health, maintain backups and patches, and define on-call ownership, runbooks, and improvement cycles.

    Outcome
    Operational dashboards, recovery evidence, incident procedures, and a continuous improvement backlog.
    Example
    Alert on failed portal requests, rehearse a database restore against agreed recovery targets, and use incident reviews and utilization data to improve runbooks and resource sizing.

In multi-cloud estates, each platform needs its own landing zone, but strategy, governance, security, and operations standards should be shared rather than reinvented per cloud.

Focusing on the five pillars

Based on the Well-Architected Framework (Microsoft Learn, opens in a new tab). CAF establishes organizational direction and platform standards; WAF guides the design and operation of each workload within them.

The pillars involve tradeoffs: no workload maximizes all five. Organizations decide how much to invest in each based on business criticality, compliance needs, and time to market. The examples use the same hypothetical customer portal as the CAF phases; numerical targets are illustrative and must be agreed with the business.

Workloadbalanced tradeoffsReliabilitySecurityCostOptimizationOperationalExcellencePerformanceEfficiency
Each pillar contributes design principles, a review checklist, and tradeoffs; the workload sits where those decisions meet.

Reliability

Keep critical user journeys available and recover within agreed limits. Set an SLO for service quality, an RTO for restoration time, and an RPO for acceptable data loss. Analyze dependency failures, remove single points of failure, and test recovery.

Example
Run the portal across availability zones, handle transient dependency failures with bounded retries, and rehearse database restoration. Check that the whole service meets its recovery targets.
Tradeoff
Extra replicas and standby regions increase cost and operational complexity; use the business impact of downtime to justify them.

Security

Protect confidentiality, integrity, and availability according to the workload’s risks. Classify data, model threats, apply least privilege and segmentation, encrypt sensitive information, and monitor for attacks. Plan how to contain and recover from incidents.

Example
Give the portal’s managed identity only the database permissions it needs. Store secrets in a vault, keep data access private, redact sensitive logs, and alert on unusual access patterns.
Tradeoff
Stronger isolation and access controls add configuration and support effort; make secure paths easy for teams to use.

Cost Optimization

Deliver business value within a sustainable budget. Model total cost, including compute, storage, network transfer, licenses, and operations. Assign cost owners, track cost per useful transaction, and tune capacity and pricing using measured demand.

Example
Stop unused development environments, rightsize an underused database after load testing, and reserve capacity only for a stable baseline. Track portal cost per successful customer request.
Tradeoff
Reservations reduce flexibility, and aggressive downsizing can hurt reliability or response time. Recheck service targets after every saving.

Operational Excellence

Make development and operations repeatable, observable, and safe. Use versioned infrastructure, automated tests and deployment, clear ownership, actionable telemetry, and incident runbooks. Turn production feedback into improvements.

Example
Deploy the portal through a pipeline with infrastructure checks and a canary release. Watch error rates and latency, roll back when thresholds are breached, and rehearse the response with the on-call team.
Tradeoff
Automation and observability need upfront work and ongoing maintenance. Start with frequent or high-risk tasks and signals that lead to action.

Performance Efficiency

Meet user expectations as demand changes. Define latency and throughput targets for critical flows, measure the full request path, load-test realistic usage, and choose scaling and data-access patterns that address measured bottlenecks.

Example
Set an illustrative target of 95% of portal requests completing within 300 ms at expected peak load. Trace slow database queries, cache suitable reads, and scale workers using queue depth.
Tradeoff
Caching introduces freshness decisions; spare capacity improves response time but costs more. Autoscaling also needs time to react to sudden demand.

In practice: learn every pillar's design principles, work through the checklists by business priority, record tradeoffs explicitly, and improve in stages with the maturity model, treating the assessment as a moving score that evolves with the workload.