Skip to main content
Post your resume and find your next job on Indeed!

Reliability Maintenance Engineer jobs in Toronto, ON

Sort by: -
    • Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-…
    • As a SRE, you will implement, measure and gather insights from Operational Level Indicators identifying areas for service improvements covering availability,…
    • View all Scotiabank jobs - Toronto jobs - Site Reliability Engineer jobs in Toronto, ON
    • Salary Search: Site Reliability Engineer (SRE) salaries in Toronto, ON
    • See popular questions & answers about Scotiabank
    • Inspects grounds, facilities, and infrastructure support systems, and their performance to determine necessity of repairs or maintenance, and conducts scheduled…
    • 3 years experience in engineering, reliability, maintenance, rotating equipment, technical sales support, or industrial applications.
    • The Engineer will work 8 or 12 hour shifts in the maintenance of plant and building equipment or plant operation during days, evenings, nights, weekends or…
    • You'll play a key role in ensuring our rolling stock systems meet technical requirements and customer expectations.
    • Customer-focused mindset with a commitment to reliability and operational excellence.
    • Work with engineering teams to improve reliability, reduce recurring…
    • Envision and implement changes that improve system reliability.
    • Review, digest, and distill complex code and technical topics to ensure clarity and…
  • View similar jobs with this employer
    • The successful candidate will combine deep expertise in SAP and ADM/AMS services with the ability to integrate multiple service towers into a cohesive…
  • View similar jobs with this employer
    • Understanding of ITIL processes, SLOs, SLIs, and reliability engineering principles.
    • Monitor application health, server performance, database availability, and…
  • View similar jobs with this employer
    • Identify areas for improvement and implement changes to enhance system reliability and performance.
    • Work closely with developers, operations teams, and other…
    • In this role, you will be responsible for architecting, testing, and executing enterprise-wide software updates and security patches across diverse platforms.
    • We are seeking a Systems Safety Specialist (Systems Assurance) to support the development, integration, and assurance of safe, reliable, and compliant systems…
    • Coach, mentor, guide, and motivate a senior team of cross-functional cloud platform engineers.
    • We’re looking for a Manager, Site Reliability Engineering to lead…
    • Providing guidance to the engineering teams on implementing safety and reliability into their application and product designs.
Get email updates for the latest Reliability Maintenance Engineer jobs in Toronto, ON

By creating a job alert, you agree to our Terms . You can change your consent settings at any time by unsubscribing or as detailed in our terms.

People also searched:

reliability engineer

Career Resources:

Site Reliability Engineer
Toronto, ON
Permanent

Job details

Pay information not provided
Permanent
On call
Toronto, ON

Full job description

Employment Status: Permanent
Schedule: 40 hours/week – 100% remote work



Job Description

We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.

Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services. This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.

You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.



Responsibilities

  • Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
  • Monitor and improve service availability, latency, performance, and overall system health.
  • Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
  • Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
  • Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
  • Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
  • Coordinate incident response and act as Incident Commander during critical production events.
  • Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
  • Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
  • Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.


Required profile

  • Significant experience in Site Reliability Engineering, Cloud Engineering, DevOps, or a similar infrastructure-focused role.
  • ???????Experience supporting complex or large-scale SaaS environments with high availability requirements.
  • Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
  • Proven experience operating and troubleshooting Kubernetes environments at scale.
  • Strong knowledge of Infrastructure as Code, particularly Terraform.
  • Experience building or maintaining CI/CD pipelines using GitLab, Jenkins, or comparable tools.
  • Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
  • Strong scripting skills using Python, Bash, or a similar language.
  • Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
  • Familiarity with Java or .NET application environments is considered an asset.
  • Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
  • Must be legally authorized to work in Canada.


What to Expect

  • Remote-first work environment within Canada.
  • ???????Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
  • Participation in a scheduled on-call rotation is required.


Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to e.henry@totemtalent.ca.

Thank you for your interest in this position; only candidates who meet our client’s requirements will be contacted.


The masculine gender is used as a neutral form.


#totemtech

Let Employers Find YouUpload Your Resume