ChallengeA San Francisco-based vertical SaaS company ($68M ARR, 2,400 enterprise customers, 340 employees, Series D — $180M total raised) serving the commercial real estate industry needed to modernize its platform from a Ruby on Rails monolith to a cloud-native microservices architecture as part of pre-IPO preparation. The platform — CRE Workflow (commercial real estate portfolio management, deal pipeline, lease administration, and financial analysis) — had been built on Rails starting in 2014 when the company had 3 engineers and zero revenue. Nine years and $68M ARR later, the original Rails monolith had become the company's primary constraint. (1) The Rails monolith — 9 years of accumulated debt. Codebase: 1.2 million lines of Ruby (Rails 5.2 — two major versions behind current Rails 7), with: 4,800 ActiveRecord models (many with 30+ associations and 1,000+ line model files), 2,200 controller actions, 380 background job types (Sidekiq), 1,400 database migrations (PostgreSQL — single database serving all tenants), and 280 service objects (inconsistently structured — some containing hundreds of lines of business logic, others being thin wrappers). Technical debt inventory: no formal API layer (frontend communicated via Rails views and ad-hoc JSON endpoints — preventing mobile app development), single PostgreSQL database (all 2,400 tenants sharing one database — creating multi-tenancy risk and preventing data residency for international customers), monolithic deployment (entire application deployed as one unit — a change to the lease administration module required deploying the deal pipeline, financial analysis, and all other modules), test suite (14,000 tests taking 2 hours 40 minutes to run — developers running partial test suites and discovering failures in CI), no feature flags (changes shipped to all customers simultaneously — enterprise customers requiring change management), and logging and audit (minimal structured logging, no audit trail for data changes — SOC 2 CC7.2 gap). (2) SOC 2 Type II failures. The company had attempted SOC 2 Type II certification 8 months prior and failed on 11 controls — all rooted in monolithic architecture: CC6.1 (logical access controls): monolithic codebase meant developers had access to all customer data — no service-level isolation. CC6.3 (access provisioning/deprovisioning): no granular RBAC — 5 roles (admin, manager, analyst, viewer, API) with broad permissions. CC7.1 (detection of unauthorized changes): no immutable audit log — data changes not tracked systematically. CC7.2 (monitoring): minimal structured logging — no ability to detect anomalous access patterns. CC8.1 (change management): monolithic deployment without feature flags — no controlled rollout capability. A1.1 (availability commitments): 99.2 percent uptime (below 99.9 percent SLA commitment to enterprise customers) — monolithic architecture meaning any module failure affected entire platform. The auditor's remediation guidance explicitly recommended "architectural decomposition to enable granular access controls, independent deployment, and comprehensive audit logging." (3) Enterprise sales blocked. The company's pipeline contained $34M ARR in enterprise opportunities requiring capabilities the monolith could not provide: data residency (European customers requiring EU data storage — single US database architecture preventing this), SSO/SAML (enterprise authentication — partially implemented but unreliable on legacy architecture), dedicated environments (largest prospects requiring isolated deployments — not feasible on shared monolith), custom SLAs (99.99 percent uptime commitments — impossible on monolithic architecture with 99.2 percent actual uptime), and SOC 2 Type II (table stakes for enterprise procurement — failed certification blocking deals). Board pressure: the board had set a 14-month IPO target. IPO underwriters had communicated that SOC 2 failure, 99.2 percent uptime, and architecture risk would impact IPO valuation by an estimated 15-25 percent ($150-250M on projected $1B+ valuation). (4) Engineering productivity crisis. 340 employees included 120 engineers — but monolithic architecture meant: deployment frequency: 2 per week (industry benchmark for SaaS: multiple daily), lead time for changes: 18 days average (from commit to production — including 2h40m test suite, manual QA, deployment coordination), incident rate: 4.2 production incidents per week (monolithic deployment creating blast radius), and developer satisfaction: 34 percent (internal survey — "the monolith" cited as primary frustration). Engineering leadership estimated that 40 percent of engineering time was spent on monolith-related overhead: waiting for tests, coordinating deployments, investigating cross-module incidents, and working around architectural limitations.
SolutionWe delivered a monolith decomposition over 44 weeks — transforming the Rails monolith into a cloud-native microservices platform while maintaining production operations for 2,400 enterprise customers and achieving SOC 2 Type II certification. (1) Domain-driven decomposition analysis. Business domain mapping: working with product and engineering leadership, we identified 8 bounded contexts within the monolith: deal pipeline (deal sourcing, qualification, financial modeling, and closing workflow), lease administration (lease abstraction, critical dates, rent schedules, and CAM reconciliation), portfolio analytics (portfolio performance, benchmarking, and reporting), financial analysis (DCF modeling, IRR calculations, cash flow projections), document management (document storage, OCR, search, and collaboration), tenant management (tenant CRM, credit analysis, and communication), integration hub (third-party data — CoStar, ARGUS, Yardi — import/export), and identity and access (authentication, authorization, audit, and tenant isolation). For each bounded context: business logic extracted from Rails models and service objects, data ownership defined (which database tables belong to which domain), inter-domain dependencies mapped (deal pipeline depends on financial analysis; lease admin depends on document management), and API contracts defined (how domains communicate — synchronous REST for queries, async events for state changes). (2) Microservices architecture. Platform: Kubernetes on AWS (us-west-2 — existing region, plus eu-west-1 for EU data residency — enabling the $34M enterprise pipeline). Each bounded context became an independent service cluster: separate PostgreSQL database per service (breaking the single-database monolith — enabling tenant isolation and data residency), independent deployment pipeline (each service deployable without affecting others), service mesh (Istio — mTLS between services, traffic management, and observability), event bus (Apache Kafka — async communication between domains), and API gateway (Kong — rate limiting, authentication, and API versioning). Multi-tenancy redesign: from shared-database to schema-per-tenant for enterprise customers (data isolation) and shared-schema with row-level security for SMB customers (cost efficiency). Data residency: tenant data routable to US or EU regions based on customer preference — enabling European enterprise sales. SOC 2 architecture: every design decision evaluated against the 11 failed controls. Audit logging: immutable event log (append-only, cryptographically chained) recording every data access and modification — CC7.1 compliant. RBAC: fine-grained role-based access with service-level permissions — CC6.1 and CC6.3 compliant. Monitoring: structured logging, distributed tracing (Jaeger), and anomaly detection — CC7.2 compliant. Change management: feature flags (LaunchDarkly), canary deployments, and automated rollback — CC8.1 compliant. (3) Strangler fig migration — service by service. Migration sequence (lowest risk to highest): Identity and access (weeks 6-10): extracted first — enabling the RBAC and audit logging infrastructure that all other services depend on. New authentication (Auth0 integration replacing custom Rails authentication), SAML/SSO (enterprise requirement), fine-grained RBAC, and comprehensive audit logging. This service alone addressed 6 of the 11 failed SOC 2 controls. Document management (weeks 8-14): relatively isolated domain — document storage, OCR, and search extracted to independent service. S3-backed storage with encryption at rest and in transit. Integration hub (weeks 12-18): third-party integrations (CoStar, ARGUS, Yardi) extracted — enabling independent scaling and failure isolation (a CoStar API outage no longer affecting the entire platform). Tenant management (weeks 14-20): CRM functionality extracted. Data migration: tenant records migrated from shared database to dedicated service database. Portfolio analytics (weeks 18-26): reporting and analytics extracted. This was the first service requiring complex data migration — analytics queries spanning multiple domains required a read-replica strategy. Financial analysis (weeks 22-30): computation-heavy service (DCF calculations, IRR, cash flow modeling) extracted — enabling independent scaling of computation resources during peak analysis periods. Lease administration (weeks 26-36): complex domain — lease abstraction, critical dates, rent schedules. 480 ActiveRecord models related to leasing extracted with full data migration. Lease admin was the highest-risk extraction due to: financial accuracy requirements (rent calculations must be exact), critical date management (missed lease dates = financial liability), and tenant relationships (lease data linked to tenant, deal, and financial domains). Deal pipeline (weeks 30-40): the core product domain — deal sourcing, qualification, and closing. Final monolith extraction completing the decomposition. (4) Data migration strategy. Each service extraction required data migration from the shared PostgreSQL database to service-specific databases: change data capture (Debezium on PostgreSQL WAL): real-time data replication from monolith database to new service databases during transition period. Dual-write verification: during parallel running, both monolith and new service processing identical operations — automated reconciliation catching any discrepancy. Progressive traffic shifting: 5 percent to 25 percent to 50 percent to 100 percent of traffic for each service — monitored at each stage. Zero-downtime migration: no maintenance windows — all migrations performed live with automatic rollback capability. Total data migrated: 4.2TB across 8 service databases (from single 4.2TB monolith database). Migration accuracy: 100 percent — zero data discrepancies detected across all 8 service migrations. (5) SOC 2 Type II certification. Re-audit commenced at week 36 (while final 2 services were still being extracted — sufficient evidence from 6 completed services). Evidence collection automated: audit logging generating continuous SOC 2 evidence, access reviews automated through RBAC system, change management documented through deployment pipeline, and monitoring dashboards demonstrating CC7.2 compliance. Result: SOC 2 Type II certified at week 42 — passing all 11 previously failed controls plus 4 additional controls that the auditor flagged as "exceeding requirements." (6) Engineering productivity transformation. CI/CD per service: test suites reduced from 2h40m (monolith) to 3-12 minutes per service. Deployment: from 2x per week to 15-25 deployments per day across services. Feature flags: controlled rollout to customer segments — enterprise customers receiving changes after SMB validation. Incident blast radius: service-level isolation — a financial analysis issue no longer affecting deal pipeline or lease administration.
OutcomeResults over 12 months post-completion. SOC 2 Type II: certified at week 42 (from failed audit 8 months prior). All 11 previously failed controls passed. Continuous compliance: automated evidence generation maintaining certification. Enterprise pipeline unlocked: $34M ARR enterprise pipeline accessible (data residency, SSO/SAML, dedicated environments, SOC 2, custom SLAs all enabled). Closed $18.4M ARR in enterprise deals within 9 months of SOC 2 certification. European expansion: EU data residency enabled — 3 enterprise customers (combined $4.2M ARR) requiring EU-only data storage. Platform reliability: uptime from 99.2 percent to 99.97 percent (monolith: 44 hours annual downtime to 2.6 hours). Production incidents from 4.2 per week to 0.6 per week (86 percent reduction). Mean time to recovery from 2.4 hours to 8 minutes (service-level isolation enabling rapid rollback). Engineering productivity: deployment frequency from 2x per week to 18x per day (average across services). Lead time from 18 days to 1.8 days (90 percent reduction). Test suite from 2h40m to 3-12 minutes per service. Developer satisfaction from 34 percent to 78 percent. Revenue impact: ARR from $68M to $92M over 12 months ($24M growth — modernization-enabled enterprise deals being the primary driver). Net revenue retention from 108 percent to 124 percent (enterprise upsells enabled by platform capabilities). IPO preparation: technology due diligence completed at week 48 — underwriters noting "material improvement in technology risk posture" and removing the 15-25 percent valuation discount. SOC 2 certification and 99.97 percent uptime meeting all underwriter requirements. Total modernization investment: $4.8M over 44 weeks (including ZTABS engagement, infrastructure migration, and internal engineering allocation). Enterprise revenue enabled: $22.6M ARR (closed $18.4M + pipeline conversion). Annualised infrastructure cost: from $2.1M (monolith — oversized instances, inefficient resource utilization) to $1.6M (microservices — right-sized, auto-scaling). Net financial impact: $4.8M investment, $22.6M new ARR enabled, $500K annual infrastructure savings, and IPO valuation preservation estimated at $150-250M.