Overview
NOTE: Hybrid 3 days a week in Toronto Office
Type: 9 month contract, 8 hours/day, 40 hours/week
Rate: $85 to $94/hr
SKILLS: Python, GitHub Actions, CI/CD, automation, Site Reliability Engineering, Dynatrace, Argo CD, PowerShell, observability
Industry: Capital Markets and Investment Management
DESCRIPTION
We are seeking an experienced Site Reliability Engineer to support applications, data pipelines, platforms, and services across Capital Markets, Investment Analytics, and Total Fund Management.
This is a hands on individual contributor role focused on Python development, operational automation, CI/CD, observability, disaster recovery readiness, and platform modernization. The successful candidate will improve health checks, deployment pipelines, telemetry, diagnostics, and operational processes while addressing a significant backlog of reliability initiatives.
The environment is highly Python focused and includes GitHub Actions, Argo CD, Dynatrace, Elastic and ELK, PowerShell, Ansible, Bash, and cron jobs. Incident volume is relatively low, allowing the successful candidate to focus primarily on proactive engineering and sustainable automation.
Capital markets experience is beneficial but not required. Strong Python, automation, CI/CD, and observability experience are the primary priorities.
RESPONSIBILITIES
• Build Automation and Scripting Solutions: Develop maintainable Python scripts, tools, and automation solutions that reduce manual effort, improve consistency, and strengthen operational efficiency.
• Improve CI/CD and Deployment Processes: Enhance build, deployment, release, rollback, and production validation processes using GitHub Actions, Argo CD, and related technologies.
• Advance Site Reliability Engineering Practices: Improve the availability, performance, resilience, supportability, and operational sustainability of applications, data pipelines, platforms, and services.
• Enhance Observability and Telemetry: Strengthen application health checks, logging, monitoring, metrics, alerts, dashboards, diagnostics, and telemetry. Support the adoption of Dynatrace and the transition from the current Elastic and ELK logging environment.
• Strengthen Incident Response and Root Cause Analysis: Investigate production issues using logs, metrics, traces, and diagnostic tools. Identify root causes and implement corrective and preventative actions.
• Improve Detection and Recovery: Develop monitoring, diagnostics, automation, and operational runbooks that reduce Mean Time to Detect and Mean Time to Restore.
• Improve Operational Readiness: Conduct production readiness reviews and ensure monitoring, runbooks, resilience testing, documentation, release processes, and support arrangements are in place before production deployment.
• Support Disaster Recovery and Resilience: Contribute to disaster recovery plans, recovery procedures, resilience initiatives, and operational testing.
• Support Production Systems: Participate in incident response, coordinate technical activities, communicate service impacts, and escalate major or cross service incidents through the appropriate channels.
• Partner Across Technology Teams: Collaborate with Product Engineering, Data Solutions, Data Platform, Architecture, Security, Technology Services, database, and quality assurance teams to identify and resolve reliability, resilience, and supportability risks.
REQUIREMENTS
• 6 or more years of experience in Site Reliability Engineering, production engineering, platform engineering, application support, or technology service delivery.
• Strong hands on Python development and scripting experience, including sound coding, testing, documentation, and source control practices.
• Experience designing automation solutions and independently selecting the appropriate scripting language, platform, or tool to solve operational problems.
• Experience developing, maintaining, or improving CI/CD pipelines, preferably using GitHub Actions and Argo CD.
• Experience with observability, application monitoring, centralized logging, telemetry, alerting, and production diagnostics. Dynatrace experience is strongly preferred.
• Experience with Elastic, ELK, PowerShell, Ansible, Bash, cron jobs, or similar scripting and automation technologies is beneficial.
• Hands on experience operating and improving highly available applications, data pipelines, platforms, APIs, message queues, or distributed services.
• Experience with incident response, root cause analysis, disaster recovery, resilience testing, capacity planning, and production readiness.
• Experience with cloud environments, infrastructure as code, DevOps, DataOps, automated testing, release automation, and production validation practices.
• Experience using Git, JIRA, Confluence, and related engineering collaboration tools.
• Strong problem solving skills with the ability to understand an unfamiliar problem space, establish an effective approach, and independently drive work through completion.
• Strong communication and collaboration skills, including the ability to explain technical risks, incidents, and service impacts in clear business language.
• Experience working with developers, delivery leads, architects, business analysts, database administrators, quality assurance professionals, security teams, and infrastructure teams.
• Knowledge of Agile, Waterfall, DevOps, ITIL, or COBIT practices.
• Practical experience using approved AI assisted engineering tools to troubleshoot applications, databases, infrastructure, code, logs, and telemetry is an asset.
• Capital markets, investment analytics, investment management, or financial services experience is beneficial but not required.
Note: As part of our hiring process, we use AI based systems to support initial applicant screening.