ChallengeA Toronto-based life insurance company (CAD $18B in in-force policies, 2.8M policyholders across Canada, offering individual life, group benefits, and wealth management, with 4,200 employees and operations in all 10 provinces) needed a cloud migration and DevOps platform to modernise infrastructure that was limiting digital transformation. The company operated 180 production applications on aging on-premises infrastructure in two Toronto-area data centres, with technology decisions constrained by 25-year-old architecture choices. Core challenges: (1) Data centre end-of-life — the company's primary data centre lease expiring in 18 months with no renewal option (the building being converted to residential). The secondary data centre serving as DR could not accommodate the full production workload. The company facing a forced infrastructure migration — not an optional modernisation project but a mandatory move with a hard deadline. 180 applications needed to move somewhere: another data centre, cloud, or a combination. The CTO choosing cloud as the strategic direction but facing the challenge of migrating 180 applications (many running on Windows Server 2012/2016 with dependencies on legacy middleware) within 18 months while maintaining 24/7 availability for policyholder services. (2) OSFI compliance in cloud — OSFI Guideline B-13 requiring the company to: maintain adequate technology and cyber risk management in cloud environments, ensure data residency for policyholder information (Canadian-hosted only), maintain business continuity capabilities, and demonstrate to OSFI examiners that cloud infrastructure met the same standards as on-premises. The company's previous cloud pilot (a non-critical HR application on Azure) had taken 6 months and generated 140 pages of OSFI compliance documentation — at that pace, migrating 180 applications with proper compliance would take years. DevOps automation was essential to make cloud migration manageable within OSFI constraints. (3) Legacy application complexity — the 180 applications spanning: a mainframe-era policy administration system (AS/400, running COBOL with RPG interfaces, managing all 2.8M policies), 40 Java applications (various versions from Java 6 to Java 17), 60 .NET applications (ranging from .NET Framework 3.5 to .NET 6), 25 database-centric applications (Oracle 12c, SQL Server 2016/2019), and 35 vendor-packaged applications (various configurations and licensing constraints). Integration: 2,400 point-to-point integrations between these applications — moving one application required understanding and maintaining all its dependencies. The company's enterprise architecture team estimating 12 months just to document all integrations, before any migration work began. (4) Provincial variation — the company operating under 10 provincial insurance regulators plus OSFI federal oversight. Applications needed to handle: different regulatory requirements by province, different tax calculations, different product features (Quebec French-language requirements), and different reporting obligations. Infrastructure changes needed to maintain provincial compliance — a deployment that broke Quebec language requirements or Ontario rate filing calculations would create regulatory exposure. (5) Change management immaturity — deployments managed through a Change Advisory Board (CAB) meeting weekly, with deployment requests submitted via email, and changes executed by a team of 6 operations engineers using manual runbooks. Average deployment lead time: 28 days (from request to production). Change failure rate: 15 percent (deployments requiring rollback or immediate fix). Weekend maintenance windows: 4 hours every other Saturday night, with capacity for 3-4 deployments per window. With 180 applications to migrate, the current change process was the bottleneck — at 3-4 changes per bi-weekly window, migration alone would consume 2+ years of deployment capacity.
SolutionWe delivered cloud migration infrastructure and DevOps platform over 16 weeks — enabling the migration of 180 applications to AWS within the 18-month deadline with OSFI-compliant automation. (1) Migration factory: automated, repeatable application migration. Assessment automation: automated discovery scanning all 180 applications — cataloguing: operating system, runtime versions, database dependencies, network connections (all 2,400 integrations mapped automatically through network flow analysis), data classification (policyholder PII requiring Canadian residency), and cloud-readiness scoring. Applications classified: "lift and shift" (85 applications suitable for direct EC2/container migration), "modernise" (55 applications requiring containerisation or re-platforming), "refactor" (25 applications requiring code changes for cloud compatibility), and "retain" (15 applications remaining on-premises including the AS/400 policy administration system — connected to cloud through secure hybrid networking). Migration pipelines: standardised migration pipelines for each category — Terraform modules for infrastructure provisioning, Ansible playbooks for application configuration, and automated testing suites for post-migration validation. Each migration following the same automated process: provision infrastructure, deploy application, run validation tests, migrate data, switch DNS, monitor for 48 hours, decommission old infrastructure. A migration taking 2-3 days per application (versus the previous 2-3 months) through automation. Wave planning: applications grouped into migration waves based on dependency analysis — applications with no cross-dependencies migrated first, building confidence and process maturity before tackling tightly coupled application clusters. 12 waves planned across 14 months, with the first 3 waves (low-risk applications) completing in weeks 8-16 of the engagement. (2) OSFI-compliant cloud platform: Canadian-hosted, regulated infrastructure. Landing zone: AWS ca-central-1 (Montreal) as primary region with ca-west-1 (Calgary) as DR — both Canadian regions ensuring data residency for all policyholder information. Landing zone architecture meeting OSFI B-13 requirements: multi-account structure (separate accounts for production, staging, development, shared services, and logging), network isolation (VPC per environment with strict security group rules), and centralised logging (all API calls, access events, and infrastructure changes logged to a tamper-evident audit account). Guardrails: AWS Control Tower with custom Service Control Policies (SCPs) preventing: resource creation outside Canadian regions, unencrypted data storage, public network access to production resources, and IAM policy modifications without approval workflow. These guardrails ensuring that even misconfigured deployments couldn't violate OSFI requirements. Hybrid connectivity: AWS Direct Connect from both Toronto data centres to ca-central-1 — the AS/400 system remaining on-premises but accessible from cloud applications through a secure, low-latency connection. Transit Gateway managing routing between cloud VPCs and on-premises networks. (3) CI/CD platform: automated, compliant deployment. Pipeline standardisation: every application receiving a CI/CD pipeline (GitHub Actions for CI, Terraform + Ansible for deployment) with mandatory stages: code scanning (SonarQube for quality, Snyk for vulnerabilities), build and unit testing, integration testing (against dependent services), OSFI compliance verification (automated checks for: data residency, encryption, access controls, and logging), staging deployment and smoke tests, approval gate (automated for low-risk changes, manual for database migrations and security configuration changes), and production deployment with automated validation. Deployment strategies: blue-green for policyholder-facing applications (zero-downtime switchover), rolling updates for internal applications, and canary releases for high-traffic applications. All strategies with automated rollback — health checks monitoring error rates, response times, and business metrics (policy lookup success, claims submission success). CAB automation: the weekly Change Advisory Board replaced with automated change management — low-risk changes (code-only deployments that passed all automated gates) proceeding automatically, medium-risk changes (configuration changes, dependency updates) requiring one approval, and high-risk changes (database migrations, security changes) requiring two approvals and scheduled windows. Change documentation automated — every deployment generating a complete change record. (4) Observability and incident response: production-grade monitoring. Monitoring: Datadog deployed across all migrated applications — APM for Java and .NET applications, infrastructure monitoring for all AWS resources, custom metrics for business operations (policy issuance rate, claims processing throughput, premium collection success). Provincial monitoring: dashboards segmented by province — ensuring that a deployment affecting Ontario policy processing was immediately visible, even if national metrics looked healthy. Incident management: PagerDuty with OSFI-aligned escalation — critical incidents (policyholder service disruption) triggering immediate engineering response plus OSFI notification preparation (OSFI Technology and Cyber Security Incident Reporting requirements specifying reporting within specific timeframes for material incidents). DR testing: automated monthly DR failover from ca-central-1 to ca-west-1 — testing application failover, data replication integrity, and recovery time. RTO of 2 hours for critical policyholder services (versus previous 12-hour manual DR process). (5) Security and compliance automation: OSFI examination readiness. Security scanning: every deployment scanned for vulnerabilities (container images, dependencies, infrastructure configuration). CIS benchmarks applied to all AWS accounts — automated remediation for configuration drift. Compliance reporting: automated OSFI examination preparation — quarterly reports showing: change management statistics (volume, failure rates, rollback rates), security posture (vulnerability counts, patching compliance, access review status), availability metrics (uptime by service, incident counts, MTTR), and cloud governance (data residency compliance, encryption status, access controls). Penetration testing: automated security testing integrated into CI/CD pipelines — OWASP checks on every deployment, quarterly comprehensive penetration testing of the cloud environment.
OutcomeResults over 18 months. Migration: 165 of 180 applications migrated to AWS within 14 months (ahead of 18-month deadline). 15 applications retained on-premises (AS/400 and tightly coupled legacy systems) connected via Direct Connect hybrid architecture. Zero policyholder service disruptions during migration — blue-green deployments ensuring seamless cutover for every application. Data centre lease: primary data centre decommissioned 4 months before lease expiry — no emergency extension needed, saving CAD $1.8M in potential holdover costs. Deployment: deployment frequency from bi-weekly (3-4 changes per window) to average 12 deployments per day across all applications. Lead time from 28 days to 3 hours for standard changes. Change failure rate from 15 percent to 2.4 percent. CAB meeting time from 4 hours per week to 30 minutes per week (reviewing only high-risk changes). OSFI compliance: OSFI examination 10 months post-migration — zero findings related to cloud infrastructure or technology risk management. OSFI examiner noting "well-documented cloud governance framework with appropriate controls." Compliance documentation generated automatically — no manual preparation needed for examination. OSFI B-13 compliance demonstrated through: automated guardrails (preventing non-compliant configurations), complete audit trails (every infrastructure change logged), and regular DR testing (monthly automated failover). Infrastructure cost: on-premises data centre costs eliminated: CAD $4.2M annually (colocation, hardware maintenance, power, cooling, facilities staff). AWS costs: CAD $2.8M annually. Net infrastructure savings: CAD $1.4M per year. Cloud cost per application: 34 percent lower than on-premises equivalent — through right-sizing, Reserved Instances, and automated environment management. Reliability: production incidents from 6.2 per month (on-premises) to 1.8 per month (cloud). MTTR from 4.5 hours to 35 minutes. DR failover from 12 hours (manual) to 1.8 hours (automated) — within the 2-hour RTO target. Availability from 99.4 percent to 99.95 percent for policyholder-facing services. Engineering productivity: operations team from 6 engineers managing manual deployments to 3 engineers managing automated platform (3 redeployed to cloud architecture and security roles). Developer deployment involvement from 8 hours per deployment to 15 minutes (pipeline management). New application deployment from 6 weeks to 2 days — standardised platform accelerating new project delivery. Financial impact: data centre elimination CAD $4.2M (gross); cloud costs CAD $2.8M; net infrastructure savings CAD $1.4M annually. Deployment efficiency CAD $2.2M (engineering time savings). Reliability improvement CAD $1.8M (reduced incident costs, faster recovery). Holdover cost avoidance CAD $1.8M (one-time). Migration acceleration CAD $3.2M (completing ahead of schedule freed engineering capacity for other projects). Total annual impact: CAD $5.4M recurring plus CAD $5.0M one-time. Engagement cost: CAD $720K. Annual platform management: CAD $195K. ROI: 7.5x first year (recurring impact only), 14.4x including one-time savings.