TI · Ciudad de México · Financial services and fintech
Platform engineering and reliability for a Canadian fintech
De un vistazo
- Industry
- Financial services and fintech
- Client geography
- Canada, national fintech
- Client size
- Scale-up, digital financial platform
- Service line
- IT outsourcing, platform engineering, DevOps, reliability
- Primary language
- English and Spanish
- Delivery site
- Mexico City (CDMX)
- Engagement duration
- 16 months, ongoing
- Team size
- 20 engineers, 2 leads, 7 senior, 9 mid, 2 junior
Perfil del cliente
The client is a Canadian fintech operating a digital financial platform, headquartered in Canada with a distributed engineering organisation. Its platform handles payments and financial transactions at scale, where availability and reliability are existential.
El reto
The fintech's platform availability stood at 97.4%, which for a payments platform represents an unacceptable amount of downtime. Its rapid growth was outpacing the reliability of its infrastructure.
The engineering organisation was consumed by incident response and had no capacity to improve the platform. Every reliability initiative in the preceding year had been abandoned partway when incident load reclaimed the team.
Deployment was slow and risky. Without progressive delivery infrastructure, every release carried full-platform risk, which made the team release infrequently, which made each release larger and riskier.
The fintech had tried to hire senior platform and reliability engineers in Canada and found the profile scarce and expensive, with its own engineers pulled toward feature work by product pressure.
Por qué Corpshore México
Corpshore Mexico proposed a two-track engagement, a run track stabilising the platform and a build track constructing the reliability and delivery infrastructure, with the build track contractually ring-fenced from incident response so it could not be consumed by firefighting.
Mexico City placed the team in the country's financial and engineering centre, on the client's clock across the full US and Canadian working day, and Corpshore's experience with regulated financial platforms and its security posture mattered for a payments environment.
La colaboración
Twenty engineers in Mexico City: two technical leads, seven senior, nine mid-level and two junior, split across a run track and a ring-fenced build track. Run coverage is 24 hours for critical incidents; the build track works standard hours and is contractually protected from incident response. Stack: Kubernetes, Terraform, AWS, Go, Python, Prometheus, Grafana.
Enfoque y metodología
Stabilise by cause, not by ticket. Incident response moved from firefighting to cause elimination, with every critical incident producing a root cause record and a preventive action.
Ring-fenced reliability build. The build track constructed observability, progressive delivery and self-service infrastructure, protected from incident load, which is the only reason it survived to deliver where the fintech's prior internal attempts had not.
Progressive delivery. Feature flags, canary deployments and automated rollback decoupled release from risk.
Documented knowledge transfer. The infrastructure and runbooks are built to a standard the client's own engineers operate.
Resultados
Core platform availability rose from 97.4% to 99.88% by month 12, reducing monthly downtime from roughly nineteen hours to under one.
Mean time to resolution on critical incidents fell from over eleven hours to 0.6 hours, and critical incident frequency fell as cause elimination took hold.
The fintech's own engineers were returned to feature work as the platform stabilised, and deployment frequency rose to weekly or better per team.
Indicadores clave
| Indicador | Inicial | Después | Cambio |
|---|---|---|---|
| Core platform availability | 97.4% | 99.88% | +2.48 pts |
| Monthly downtime | ≈ 19 hours | < 1 hour | -95% |
| Mean time to resolution, critical | 11 h 05 m | 36 min | -95% |
| Critical incidents per month | 12 | 3 | -75% |
| Deployment frequency | Infrequent | Weekly+ | Structural change |
| Change failure rate | 21% | 7% | -67% |
| Reliability build initiatives completed | 0 of prior attempts | Delivered | New capability |
Every reliability programme we had run was eaten by incidents. Ring-fencing the build team contractually is the only reason this one survived. On a payments platform, that reliability is the business.
Valor duradero
The observability and progressive delivery infrastructure are client-owned and operated by the fintech's engineers. The ring-fenced two-track model has become the client's template. Corpshore Mexico has extended the engagement to security engineering.
Temas
Corpshore Mexico