ChallengeA London-based challenger bank (4.2M customers, FCA/PRA authorised, operating current accounts, savings, loans, and investments, with 680 employees and GBP 8.4B in deposits, growing customer base 35 percent YoY) needed a DevOps platform transformation to scale beyond its startup-era infrastructure. The bank had built its initial platform on a monolithic architecture during its founding phase — pragmatic for launch speed but now creating scaling, reliability, and compliance challenges as the bank approached the PRA's threshold for enhanced supervision. Core challenges: (1) Scaling ceiling — the bank's monolithic core banking application (Node.js, running on AWS ECS) hitting performance limits during peak usage. Morning login surge (8-9 AM, 2.8M app opens) causing response time degradation from 200ms to 1,200ms. Salary payment processing (monthly payroll credits — 1.8M transactions concentrated in 2-3 day windows) causing queue backlogs that delayed customer notifications by 4-6 hours. The engineering team had spent 6 months attempting to optimise the monolith — achieving 30 percent improvement but still hitting architectural limits. Vertical scaling (larger instances) providing diminishing returns. The bank needed to decompose into microservices — but doing so safely while maintaining 24/7 banking service for 4.2M customers. (2) Deployment risk — the monolith deployed as a single unit — every change to any component required deploying the entire application. With 85 engineers contributing to the codebase, deployment conflicts were constant. The bank deploying to production once per week in a 4-hour Sunday night maintenance window — the only time the risk team was comfortable with a full deployment. Deployment failure rate: 22 percent of weekly deployments requiring rollback. Each rollback taking 45 minutes (full application rollback) during which customers experienced degraded service. The weekly deployment cadence meaning that bug fixes took an average of 10 days to reach customers — including critical security patches. (3) PRA operational resilience — the PRA's operational resilience framework (PS6/21, PS21/3) requiring the bank to: identify Important Business Services, set impact tolerances (maximum tolerable disruption), and test resilience through scenario testing. The bank's IBS included: payment processing, account access, and customer support — each with impact tolerances measured in hours. The current infrastructure couldn't demonstrate resilience: no automated failover, no chaos engineering, and DR testing done manually once per year (and failing the last test — failover taking 6 hours versus the 2-hour impact tolerance for payment processing). The PRA supervisor expressing "concern" about the bank's technology resilience during the most recent periodic summary meeting. (4) FCA change management — the FCA expecting regulated firms to maintain documented change management processes. The bank's change process was semi-formal: Jira tickets, Slack approvals, and deployment scripts maintained by individual engineers. No separation of duties — the same engineer who wrote code could deploy it to production. No automated compliance verification — configuration changes (database migrations, feature flags, rate limits) deployed without formal approval. The bank's internal audit flagging change management as a "significant control weakness." (5) Cloud cost trajectory — AWS costs growing 45 percent annually, reaching GBP 3.8M — significantly outpacing customer and transaction growth (35 percent). The bank running 340 ECS tasks across 12 services, plus 28 RDS instances, 15 ElastiCache clusters, and 180 Lambda functions. No cost attribution — the finance team receiving a single AWS bill with no visibility into which services consumed which resources. Engineering teams provisioning "safe" instance sizes (always the next size up from what was needed) because there was no automated right-sizing.
SolutionWe delivered a DevOps platform transformation over 14 weeks — decomposing the monolith into microservices, building production-grade CI/CD, and establishing PRA-compliant operational resilience. (1) Microservice decomposition: strangler fig migration. Bounded context identification: working with the engineering team to identify 18 bounded contexts within the monolith — payments, accounts, identity, lending, savings, investments, notifications, customer support, fraud, compliance, reporting, cards, direct debits, standing orders, fees, referrals, onboarding, and analytics. Migration strategy: strangler fig pattern — new microservices deployed alongside the monolith, with traffic gradually shifted through an API gateway. No big-bang migration — the monolith continued serving traffic for domains that hadn't yet been extracted. Priority: payments extracted first (highest scale requirement and most critical IBS), then accounts, then identity — the three services representing 70 percent of load. Each extraction taking 2-3 weeks, with the microservice running in parallel with the monolith for 1 week of shadow traffic validation before cutover. Kubernetes platform: EKS cluster with: namespace-per-service isolation, Istio service mesh for inter-service communication (mTLS, circuit breaking, observability), horizontal pod autoscaling tuned per service (payments scaling to 200 pods during salary processing, accounts scaling during morning login surge), and cluster autoscaler managing node capacity. (2) CI/CD platform: fast, safe, compliant deployments. Pipeline architecture: each microservice independently deployable through standardised pipelines — GitHub Actions for CI (build, test, security scan), ArgoCD for GitOps-based Kubernetes deployment. Deployment frequency: from weekly (entire monolith) to multiple times daily per service. Developers merging to main and having changes in production within 30 minutes (including automated testing, security scanning, and staged rollout). Deployment strategies: canary releases for critical services (payments, accounts) — 5 percent traffic to new version, automated health checks over 15 minutes, progressive rollout to 100 percent if healthy. Automated rollback if error rate exceeds 0.1 percent or p99 latency exceeds 500ms. Feature flags: LaunchDarkly integration enabling feature rollout independent of deployment — engineers deploying code to production with features disabled, then enabling progressively (staff first, then 1 percent, then 10 percent, then 100 percent). FCA-compliant change management: every production change automatically documented — who, what, when, why, test results, approvals, and rollout status. Separation of duties enforced by pipeline — code review required from a different engineer, deployment approval from a different person than the author. Audit trail queryable by the compliance team and exportable for FCA examination. (3) Operational resilience: PRA-compliant resilience engineering. Impact tolerance monitoring: real-time monitoring of each IBS against its impact tolerance — payment processing monitored against 2-hour impact tolerance, account access against 4-hour tolerance, customer support against 8-hour tolerance. Dashboard visible to the CTO, CRO, and PRA supervisor. Chaos engineering: automated resilience testing using Gremlin — monthly testing of failure scenarios: availability zone failure, database failover, third-party payment scheme outage, and DNS failure. Each test measuring whether IBS remained within impact tolerances. Disaster recovery: automated DR to a second AWS region (eu-west-2 London to eu-west-1 Ireland) — RTO (Recovery Time Objective) of 15 minutes for payment processing, RPO (Recovery Point Objective) of zero for transaction data (synchronous replication). DR tested monthly through automated failover drills. Incident management: PagerDuty-based incident response with automated escalation — severity levels aligned with IBS impact tolerances. Automated communication to stakeholders (including regulatory notification triggers for incidents exceeding impact tolerances). (4) Observability: full-stack monitoring and insight. Monitoring: Datadog as the unified observability platform — application performance monitoring (APM) with distributed tracing across all 18 microservices, infrastructure monitoring, log management, and synthetic monitoring for customer-facing flows. Business metrics: real-time monitoring of business metrics — payment success rate, account opening completion rate, app response time percentiles, and transaction processing latency. These metrics reported alongside technical metrics because PRA supervision assessed business impact, not just system availability. SLOs and error budgets: each service having defined SLOs (Service Level Objectives) — payment processing at 99.99 percent success rate, account access at 99.95 percent availability. Error budgets managed formally — teams with remaining error budget could deploy more aggressively, while teams who had exhausted their budget focused on reliability improvement. Alerting: alert noise reduced from 220 weekly alerts (mostly ignored) to 30 actionable alerts — intelligent grouping, deduplication, and severity-based routing ensuring on-call engineers responded only to genuine issues. (5) Cost optimisation: cloud efficiency at scale. Cost attribution: every AWS resource tagged with service owner, environment, and cost centre — monthly cost reports by service enabling engineering teams to understand and optimise their infrastructure costs. Right-sizing: automated analysis identifying GBP 1.1M in oversized resources — RDS instances downgraded, ECS tasks right-sized, and Lambda memory configurations optimised. Savings Plans: 1-year Compute Savings Plans for stable workloads — achieving 38 percent discount on committed compute. Spot instances for non-critical workloads (batch processing, analytics, testing) — 70 percent savings versus on-demand. Environment lifecycle: non-production environments automatically scaled down outside working hours — preview environments for pull requests created and destroyed automatically, staging environments running only during business hours.