Reliability Maintenance Engineer jobs in Toronto, ON
Site Reliability Engineer
Easily applyTotem RecrutementToronto, ON- Permanent
- On call
- Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-…
- ScotiabankToronto, ON
- As a SRE, you will implement, measure and gather insights from Operational Level Indicators identifying areas for service improvements covering availability,…
- View all Scotiabank jobs - Toronto jobs - Site Reliability Engineer jobs in Toronto, ON
- Salary Search: Site Reliability Engineer (SRE) salaries in Toronto, ON
- See popular questions & answers about Scotiabank
- BMO Financial GroupToronto, ON
- $45,000–$100,000 a year
- Inspects grounds, facilities, and infrastructure support systems, and their performance to determine necessity of repairs or maintenance, and conducts scheduled…
- SKFScarborough, ON
- $82,000–$92,000 a year
- Permanent
- 3 years experience in engineering, reliability, maintenance, rotating equipment, technical sales support, or industrial applications.
- View all SKF jobs - Scarborough jobs - Application Developer jobs in Scarborough, ON
- Salary Search: Application Engineer salaries in Scarborough, ON
- See popular questions & answers about SKF
- Michael Garron HospitalToronto, ON M4C 3E7
- Full-time
- Weekends as needed +3
- The Engineer will work 8 or 12 hour shifts in the maintenance of plant and building equipment or plant operation during days, evenings, nights, weekends or…
- ALSTOMToronto, ON
- $92,000–$121,000 a year
- Full-time
- You'll play a key role in ensuring our rolling stock systems meet technical requirements and customer expectations.
- View all ALSTOM jobs - Toronto jobs - Senior Engineer jobs in Toronto, ON
- Salary Search: Senior Requirements Engineer salaries in Toronto, ON
- See popular questions & answers about ALSTOM
- Tata Consultancy ServicesToronto, ON
- On call
- Customer-focused mindset with a commitment to reliability and operational excellence.
- Work with engineering teams to improve reliability, reduce recurring…
- View all Tata Consultancy Services jobs - Toronto jobs
- Salary Search: Site Reliability Engineer salaries in Toronto, ON
- See popular questions & answers about Tata Consultancy Services
- TeslaRichmond Hill, ON
- $167,200–$252,400 a year
- Envision and implement changes that improve system reliability.
- Review, digest, and distill complex code and technical topics to ensure clarity and…
- View all Tesla jobs - Richmond Hill jobs - Site Reliability Engineer jobs in Richmond Hill, ON
- Salary Search: Staff Site Reliability Engineer, Energy Software salaries
- See popular questions & answers about Tesla
View similar jobs with this employerCapgeminiMississauga, ON- $114,847–$269,436 a year
- Permanent
- The successful candidate will combine deep expertise in SAP and ADM/AMS services with the ability to integrate multiple service towers into a cohesive…
View similar jobs with this employerCapgeminiMississauga, ON- $114,847–$269,436 a year
- Permanent
- The successful candidate will combine deep expertise in SAP and ADM/AMS services with the ability to integrate multiple service towers into a cohesive…
View similar jobs with this employerCapgeminiToronto, ON- $80,000–$90,000 a year
- Permanent
- On call
- Understanding of ITIL processes, SLOs, SLIs, and reliability engineering principles.
- Monitor application health, server performance, database availability, and…
View similar jobs with this employerCapgeminiToronto, ON- $79,000–$101,000 a year
- Permanent
- Identify areas for improvement and implement changes to enhance system reliability and performance.
- Work closely with developers, operations teams, and other…
- View all Capgemini jobs - Toronto jobs - Site Reliability Engineer jobs in Toronto, ON
- Salary Search: Site Reliability Engineer - OpenShift salaries in Toronto, ON
- See popular questions & answers about Capgemini
View similar jobs with this employerTMX Group LimitedToronto, ON- $110–$120 an hour
- Full-time
- In this role, you will be responsible for architecting, testing, and executing enterprise-wide software updates and security patches across diverse platforms.
- SYSTRAToronto, ON
- Permanent
- We are seeking a Systems Safety Specialist (Systems Assurance) to support the development, integration, and assurance of safe, reliable, and compliant systems…
- View all SYSTRA jobs - Toronto jobs - Safety Engineer jobs in Toronto, ON
- Salary Search: Systems Safety Engineer salaries in Toronto, ON
- See popular questions & answers about SYSTRA
Manager, Site Reliability Engineering
Easily applyLightspeed Commerce, Inc.Toronto, ON- $150,000–$185,000 a year
- Full-time
- On call
- Coach, mentor, guide, and motivate a senior team of cross-functional cloud platform engineers.
- We’re looking for a Manager, Site Reliability Engineering to lead…
- Hitachi RailToronto, ON
- Full-time
- Flextime
- Providing guidance to the engineering teams on implementing safety and reliability into their application and product designs.
- View all Hitachi Rail jobs - Toronto jobs - Senior System Engineer jobs in Toronto, ON
- Salary Search: Senior System RAMS Engineer salaries in Toronto, ON
- See popular questions & answers about Hitachi Rail
By creating a job alert, you agree to our Terms . You can change your consent settings at any time by unsubscribing or as detailed in our terms.
People also searched:
Career Resources:
Site Reliability Engineer
Job details
Full job description
Employment Status: Permanent
Schedule: 40 hours/week – 100% remote work
Job Description
We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.
Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services. This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.
You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.
Responsibilities
- Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
- Monitor and improve service availability, latency, performance, and overall system health.
- Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
- Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
- Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
- Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
- Coordinate incident response and act as Incident Commander during critical production events.
- Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
- Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
- Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.
Required profile
- Significant experience in Site Reliability Engineering, Cloud Engineering, DevOps, or a similar infrastructure-focused role.
- ???????Experience supporting complex or large-scale SaaS environments with high availability requirements.
- Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
- Proven experience operating and troubleshooting Kubernetes environments at scale.
- Strong knowledge of Infrastructure as Code, particularly Terraform.
- Experience building or maintaining CI/CD pipelines using GitLab, Jenkins, or comparable tools.
- Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
- Strong scripting skills using Python, Bash, or a similar language.
- Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
- Familiarity with Java or .NET application environments is considered an asset.
- Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
- Must be legally authorized to work in Canada.
What to Expect
- Remote-first work environment within Canada.
- ???????Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
- Participation in a scheduled on-call rotation is required.
Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to e.henry@totemtalent.ca.
Thank you for your interest in this position; only candidates who meet our client’s requirements will be contacted.
The masculine gender is used as a neutral form.
#totemtech