ChallengeA Stockholm-based fintech (SEK 4.2B revenue, providing business payment and expense management solutions to 38,000 companies across the Nordics, DACH, and UK, processing SEK 180B in annual payment volume, with 680 employees including 240 engineers organised into 32 squads following a Spotify-inspired organisational model, Finansinspektionen-licensed as a payment institution) needed an internal developer platform to scale engineering velocity. The company had grown from 40 to 240 engineers in 3 years while deployment velocity had actually decreased — the classic scaling problem where adding engineers slows delivery rather than accelerating it. Core challenges: (1) Squad autonomy blocked by infrastructure dependency — the 32 squads were designed to be autonomous (owning their feature area end-to-end), but infrastructure was centralised in a 4-person platform team. Every squad needing: a new database (3-day ticket), a Kafka topic (2-day ticket), a new service deployed (5-day ticket), or production access for debugging (1-day ticket with security review). The 4-person platform team processing 45 infrastructure requests per week — a 6-day average queue. Squads waiting for infrastructure rather than building features. The platform team working 60-hour weeks, burning out, and still being the bottleneck. Squads starting to work around the platform team — deploying rogue infrastructure that didn't meet security or compliance standards. (2) Deployment inconsistency — 32 squads had developed 32 different deployment approaches. Some squads had built sophisticated CI/CD pipelines (the payments squad deploying 8 times daily with canary releases). Others were still deploying through manual kubectl commands. Average deployment frequency across all squads: 1.2 per day. But this average masked enormous variation — the top 5 squads averaging 6 deployments per day while the bottom 10 squads averaged 0.3 (once every 3-4 days). Inconsistent deployment practices causing: different reliability levels across services (squads with manual deployments having 4x higher failure rates), inconsistent monitoring (some squads using Datadog, others Grafana, others nothing), and unpredictable incident response (no standardised alerting or on-call practices). (3) Finansinspektionen compliance at scale — as a licensed payment institution, the company subject to Finansinspektionen IT governance requirements: documented change management, security controls, and incident reporting. With 32 squads deploying independently, ensuring every deployment met compliance requirements was impossible — the compliance team couldn't review 40+ daily deployments. Three approaches being debated: (a) centralise all deployments through the compliance team (killing squad autonomy), (b) trust squads to self-comply (creating regulatory risk), or (c) automate compliance into the platform (the right answer, but requiring significant investment). A recent Finansinspektionen examination had noted: "inconsistent change management practices across engineering teams — recommend formalising deployment controls while maintaining development agility." (4) Developer experience degradation — despite the Spotify-model aspiration of autonomous squads, developer experience was poor. New engineers taking 4 weeks to make their first production deployment (environment setup, learning team-specific tooling, understanding team-specific deployment process). Cognitive load: each squad having its own CI/CD configuration, monitoring setup, and operational procedures meant engineers moving between squads needed to learn entirely new tooling. Development environment: engineers running 15-20 Docker containers locally to simulate the production environment — consuming 32GB RAM and making laptops unusable for other tasks. Local environment setup taking 2-3 days with frequent failures. (5) Observability fragmentation — the 32 squads using different monitoring approaches: 12 using Datadog, 8 using Grafana/Prometheus, 4 using CloudWatch, and 8 using ad hoc approaches (or nothing). No unified view of system health — an incident in one squad's service affecting another squad's service, but without cross-squad visibility, the root cause took hours to identify. Mean time to detect (MTTD): 28 minutes for squads with monitoring, unknown for squads without. Mean time to resolve (MTTR): 2.4 hours average — dominated by time spent identifying which service and which squad was responsible.
SolutionWe built an internal developer platform over 10 weeks — enabling squad autonomy through self-service infrastructure while ensuring Finansinspektionen compliance and operational consistency. (1) Self-service infrastructure: squads unblocked from platform team dependency. Service catalogue: Backstage-based developer portal providing self-service provisioning — squads creating new services, databases, message queues, and storage through a catalogue with pre-approved, compliant templates. "Create a new payment processing microservice" → service scaffold, CI/CD pipeline, database, monitoring, alerting, and documentation generated in 15 minutes (versus 5-day ticket). Golden paths: opinionated but flexible service templates — "the right way to build a service at this company" encoded in templates that included: service framework (Kotlin/Spring Boot for backend, React/Next.js for frontend), CI/CD pipeline with all compliance gates, Kubernetes deployment configuration with sensible defaults, monitoring and alerting with service-level objectives, and documentation template. Squads could customise within the template framework but couldn't deviate from security and compliance requirements. Database self-service: squads provisioning PostgreSQL, Redis, or Kafka through the platform — with appropriate sizing, backup configuration, and access controls automatically configured. Database provisioning from 3 days to 5 minutes. Environment management: each squad having automated access to: development namespace (persistent, shared within squad), staging namespace (automated deployment on PR merge), preview environments (per pull request, automatically created and destroyed), and production namespace (deployed through CI/CD pipeline with compliance gates). (2) Standardised CI/CD: consistent, compliant, fast deployment. Pipeline templates: standardised CI/CD pipeline templates providing: build (Docker multi-stage), test (unit, integration, contract tests), security scan (Trivy for images, Snyk for dependencies, Semgrep for code), compliance verification (automated Finansinspektionen policy checks), and deployment (canary release with automated rollback). Pipeline customisation: squads adding squad-specific test stages and deployment configurations within the standardised framework — maintaining consistency for security and compliance while allowing flexibility for business logic testing. Deployment guardrails: pipelines enforcing: code review (minimum 1 reviewer from the squad), security scan passing (no critical/high vulnerabilities), test coverage above squad-defined threshold (minimum 70 percent), and compliance metadata (change description, risk classification, rollback plan — auto-generated from PR description and code diff). Progressive delivery: all squads using the same progressive delivery approach — canary releases with automated health checks. New version receiving 5 percent of traffic, monitoring for 10 minutes, then progressive rollout. Automated rollback if error rate exceeds squad-defined SLO. (3) Finansinspektionen compliance automation: regulated at scale. Policy-as-code: Finansinspektionen change management requirements encoded in OPA (Open Policy Agent) policies evaluated in every deployment pipeline — ensuring: every production change has a documented description and risk classification, security scans completed within the past 24 hours, code reviewed by at least one other engineer, and rollback plan exists (automated for standard deployments). Compliance dashboard: real-time view of compliance status across all 32 squads — showing: deployment frequency, change failure rate, security scan compliance, and incident response metrics. Compliance team having visibility without being a bottleneck. Audit trail: every deployment automatically generating a compliance record — who deployed what, when, what tests passed, what security scans ran, who approved, and what the rollback procedure was. Records stored immutably for Finansinspektionen examination. Incident classification: automated incident severity classification meeting Finansinspektionen reporting thresholds — material incidents automatically generating notification drafts for the compliance team. (4) Unified observability: one view across 32 squads. Datadog standardisation: all 32 squads migrated to Datadog (retiring Grafana, CloudWatch, and ad hoc solutions) — with standardised dashboards, alerting, and APM. Not restricting squads from adding custom monitoring but ensuring a baseline of observability for every service. Service map: automatic service dependency map showing all services across all squads — when an incident occurred, immediately visible which upstream and downstream services were affected. Cross-squad incidents identified in minutes rather than hours. Standardised SLOs: every service having defined SLOs (latency, error rate, availability) with error budgets — squads consuming error budget aware of their reliability trajectory. Squads with exhausted error budgets pausing feature work for reliability improvement. On-call standardisation: every squad having defined on-call rotation with PagerDuty — standardised escalation policies, runbooks accessible through the developer portal, and post-incident review process. MTTD improved from 28 minutes to 4 minutes through standardised alerting. (5) Developer experience: fast, enjoyable, productive. Remote development environments: cloud-based development environments (Gitpod) replacing local Docker setups — engineers accessing a complete development environment in their browser within 2 minutes, with all dependencies pre-configured and production-equivalent service stubs available. Laptop RAM consumption for development from 32GB to browser tab. Onboarding automation: new engineer onboarding automated — Day 1: accounts provisioned, development environment ready, onboarding documentation in Backstage. Day 2: first code change deployed to preview environment. Day 3: first pull request submitted. First production deployment: within first week (from first month). Inner source: squads able to contribute to other squads' services through the standardised platform — same tooling, same deployment process, same monitoring everywhere. Cross-squad collaboration friction eliminated.