Daily digest

Fresh Canadian jobs in your inbox.

One email a day. Unsubscribe anytime.

Want alerts for jobs like this one?

Diverse Lynx

Systems Reliability Engineer

Diverse Lynx
Posted 6 days ago
Toronto, Ontario, CA

Top Skills Required

Service Level IndicatorsFinancial ServicesDynatrace +19 more

Description

Job Title: Systems Reliability Engineer

Location: Toronto, ON

Rate: CAD125K/Yr


Job Description:

We are seeking a Systems Reliability Engineer to join the Technical Operations team supporting our Embedded Finance (EmFi) platform. In this role, you will own the reliability and resiliency of a large-scale enterprise platform, ensuring that our services remain highly available, performant, and secure. You will design and implement monitoring and alerting frameworks, lead incident response and drive the root cause analysis (RCA) process to continuously improve platform stability. This is an opportunity to be a Partner in Possibility — helping our clients deliver financial services experiences that are essential to everyday life.


What you will do:

  • Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity.
  • Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform.
  • Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health.
  • Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams.
  • Own and advance the root cause analysis (RCA) process — investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence.
  • Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting.
  • Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency.
  • Collaborate with engineering, product, and risk stakeholders to embed reliability best practices into the platform lifecycle.

What you will need to have:

  • Hands-on experience with monitoring, observability, and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog.
  • Proven experience operating and supporting a large-scale enterprise platform environment.
  • Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes.
  • Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices.
  • Experience with ticketing and incident management workflows (e.g., ServiceNow).
  • Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams.

What would be great to have:

  • Experience in financial services, payments, or embedded finance environments.
  • Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash).
  • Familiarity with cloud platforms, containerization, and CI/CD pipelines.
  • Experience defining and managing SLOs, SLIs and error budgets











Disclaimer: Diverse Lynx LLC is an Equal Opportunity Employer. All applicants and employees are evaluated without discrimination, based solely on their qualifications, ability, competence and performance. This email and its attachments may contain confidential or proprietary information and is intended only for the recipient(s). If you received this message in error, please disregard it and notify the sender. If you no longer wish to receive our communications, you may unsubscribe at any time.
Security Notice: Our official website is www.diverselynx.com We do not operate or authorize any other websites representing Diverse Lynx LLC.

Job Details

Job Type
full time
Salary Range
$125,000 - $125,000

Required Skills

Service Level IndicatorsFinancial ServicesDynatraceContainerizationGoRoot Cause AnalysisError BudgetsCloud PlatformsPythonBashServiceNowSplunkCommunicationDatadogCI/CDReliability EngineeringIncident ResponseGrafanaMonitoring and ObservabilityCollaborationLarge-Scale Enterprise Platform OperationsService Level Objectives

Daily digest

Fresh Canadian jobs in your inbox.

One email a day. Unsubscribe anytime.