ChallengeAn Amsterdam-based payment processing company (EUR 42B annual transaction volume, serving 28,000 merchants across 15 European countries, PSD2-licensed, 320 employees, processing peak loads of 4,200 transactions per second during Black Friday) needed platform engineering to scale infrastructure that was reaching its architectural limits. The company had grown from Dutch-only to pan-European in 3 years, but infrastructure built for a single-country payment processor was buckling under multi-country regulatory and performance requirements. Core challenges: (1) Scalability ceiling — the company's payment processing engine running on AWS ECS with a PostgreSQL database cluster. The architecture handled normal load (average 800 TPS) adequately but degraded during peaks. Black Friday 2024: transaction processing latency increased from 180ms average to 2,400ms during peak hours, causing 3.2 percent of transactions to timeout — representing EUR 12M in failed payments. Merchants receiving real-time payment processing SLA guarantees (99.95 percent success rate, sub-500ms latency) — the Black Friday degradation breaching SLA for 340 merchants and triggering EUR 2.8M in SLA credit obligations. The database being the bottleneck — PostgreSQL handling both transactional writes and analytical queries (merchant reporting, reconciliation) on the same cluster, with analytical queries degrading transaction processing during peak periods. (2) Multi-country deployment complexity — operating across 15 European countries meant: payment scheme integrations (iDEAL in Netherlands, Bancontact in Belgium, SEPA across EU, Sofort in Germany, Cartes Bancaires in France, Bizum in Spain) each requiring country-specific configuration; different regulatory requirements per country; and different data residency preferences. All 15 countries running on a single AWS deployment in eu-west-1 (Ireland) — some merchants (particularly German and French) expressing concern about data not being hosted in their country. The company needing multi-region deployment but lacking the DevOps infrastructure to manage it. (3) PCI DSS compliance burden — PCI DSS Level 1 (processing over 6 million transactions annually) requiring: quarterly vulnerability scans, annual penetration testing, continuous security monitoring, and documented change management for the cardholder data environment (CDE). The company's CDE was not cleanly separated from non-CDE infrastructure — security scans flagging "scope creep" where non-payment systems shared network segments with payment processing, increasing audit scope and complexity. Every infrastructure change in the CDE requiring manual security review — adding 3-5 days to deployment lead time for payment-critical components. (4) Deployment velocity — the payment engine deploying once per week through a 6-hour Saturday maintenance window. Each deployment involving: manual database migration execution, configuration updates across 15 country-specific settings, payment scheme connectivity verification for all active schemes, and settlement reconciliation checks. A failed deployment in Month 8 had caused a 4-hour payment processing outage — the rollback taking longer than expected because database migrations were not automatically reversible. The engineering team wanting to deploy daily but constrained by the manual deployment process and the risk of payment processing disruption. (5) DORA compliance — the EU's Digital Operational Resilience Act (effective January 2025) requiring financial entities to: maintain ICT risk management frameworks, report major ICT incidents within specific timeframes, conduct digital operational resilience testing, and manage third-party ICT risk. The company's current infrastructure management was insufficient for DORA — no formal ICT risk framework, no automated incident classification, and no documented resilience testing programme.
SolutionWe delivered platform engineering over 12 weeks — rebuilding the payment processing infrastructure for European scale with PCI DSS and DORA compliance automated. (1) Microservice payment architecture: scale beyond the monolith. Service decomposition: the monolithic payment engine decomposed into 12 services — payment ingestion, routing, scheme integration (one service per payment scheme family), fraud screening, authorisation, settlement, reconciliation, merchant notification, reporting, configuration management, and monitoring. Each independently deployable, scalable, and maintainable. Event-driven architecture: Apache Kafka as the backbone — payment events flowing through topics that decoupled services. Payment ingestion publishing to Kafka, with routing, fraud, and authorisation services consuming independently. This eliminated the database bottleneck — transactional processing and analytical reporting running on separate data stores (PostgreSQL for transactions, ClickHouse for analytics). Kubernetes platform: EKS clusters in multiple AWS regions — eu-west-1 (Ireland) as primary, eu-central-1 (Frankfurt) for German/Austrian merchants, and eu-west-3 (Paris) for French merchants. Istio service mesh handling inter-service communication with mTLS, traffic routing, and circuit breaking. Auto-scaling: each service scaling independently based on its specific load pattern — the payment ingestion service scaling to handle 6,000 TPS during peaks (50 percent above previous peak), while settlement (batch-oriented) scaled on a different schedule. Scaling decisions automated through KEDA (Kubernetes Event-Driven Autoscaler) responding to Kafka consumer lag. (2) PCI DSS-compliant CI/CD: secure, fast deployment. CDE isolation: payment processing infrastructure (CDE) cleanly separated into dedicated Kubernetes namespaces with network policies preventing any communication from non-CDE workloads. Separate CI/CD pipelines for CDE and non-CDE — CDE pipelines having additional security gates but non-CDE pipelines freed from PCI scope constraints. Pipeline security: CDE deployments including: container image scanning (Trivy), dependency vulnerability checking (Snyk), infrastructure configuration scanning (Checkov for Terraform), and automated PCI compliance verification (custom policy checks ensuring: encryption at rest and in transit, no public network access, logging enabled, and access controls correct). Zero-downtime deployment: canary releases for payment processing — new versions receiving 2 percent of traffic with automated health checks monitoring: transaction success rate, latency percentiles, error rates by payment scheme, and settlement accuracy. Progressive rollout to 100 percent over 30 minutes if all checks pass. Automated rollback within 15 seconds if any metric breaches threshold. Deployment frequency: from weekly 6-hour maintenance window to daily deployments during business hours — zero maintenance windows needed because canary releases ensured no customer impact. (3) Multi-region European deployment: data residency and latency optimisation. Region strategy: payment processing traffic routed based on merchant country — Dutch merchants to eu-west-1, German/Austrian to eu-central-1, French to eu-west-3. Data residency: merchant and transaction data stored in the region matching the merchant's country, with cross-region replication only for operational redundancy (not data locality). Terraform modules: infrastructure defined once, deployed to multiple regions through parameterised Terraform modules — ensuring consistency across regions while allowing country-specific configuration (payment scheme endpoints, regulatory settings). Cross-region failover: automated failover between regions — if eu-central-1 became unavailable, German merchant traffic automatically routed to eu-west-1 within 30 seconds, maintaining service within DORA impact tolerances. (4) Observability: payment-grade monitoring. Transaction monitoring: every transaction tracked end-to-end through distributed tracing (Datadog APM) — from merchant API request through routing, fraud screening, scheme authorisation, and settlement. Transaction-level SLA monitoring: real-time dashboards showing success rate and latency per merchant, per payment scheme, and per region. Scheme health: individual payment scheme integrations monitored for: availability, response time, error rates, and settlement status. Scheme degradation detected automatically — if iDEAL response times increased 50 percent, the operations team alerted immediately with diagnostic information. Business metrics: real-time monitoring of: transaction volume, success rates, average transaction value, merchant-level performance, and settlement reconciliation status. Finance team having real-time revenue visibility rather than waiting for end-of-day reports. Alert intelligence: 180 weekly alerts reduced to 22 actionable alerts through: noise reduction, intelligent grouping, and business-impact-based severity classification. On-call engineers receiving context-rich alerts with runbook links rather than raw metric threshold breaches. (5) DORA compliance automation: digital operational resilience. ICT risk framework: documented ICT risk management framework meeting DORA Article 5-16 requirements — risk identification, protection measures, detection capabilities, response procedures, and recovery plans. All automated and continuously validated. Incident classification: automated incident classification meeting DORA reporting requirements — incidents automatically assessed for: impact on financial services, number of affected clients, cross-border impact, and data integrity. Incidents meeting DORA materiality thresholds automatically generating notification drafts for DNB. Resilience testing: quarterly TLPT (Threat-Led Penetration Testing) programme using threat intelligence to design realistic attack scenarios. Monthly chaos engineering testing: region failure, payment scheme outage, database failover, and Kafka cluster failure — all results documented for DORA examination. Third-party risk: automated monitoring of cloud provider (AWS) service health, CDN availability, and payment scheme operational status — meeting DORA third-party ICT risk management requirements.