Source Job

Canada

  • Deliver customer excellence and meet all SLAs.
  • Deploy, upgrade, and support applications, services, and operating systems.
  • Troubleshoot system performance and application health issues.

Troubleshooting Linux Windows PowerShell Cloud Platforms

20 jobs similar to Co-op/Intern Site Reliability Engineer

Jobs ranked by similarity.

US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

$4,538–$5,772/mo
Poland

  • Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.

XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

Europe

  • Support cloud-hosted application testing, implementation, maintenance, and validation.
  • Monitor and troubleshoot Windows, Linux, server, network, and application environments.
  • Assist with incident response, documentation, and operational process improvements.

Applied Systems builds cloud software and AI-powered solutions that reinvent insurance technology for agencies and brokers worldwide. With over 40 years of experience, the company fosters a people-first culture built on trust, inclusion, and growth, supporting employees to deliver their best work.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

Global

  • Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
  • Monitor production environments to ensure high availability, reliability, and performance.
  • Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.

Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.

Canada

  • Manage team, career development, project prioritization, and performance review.
  • Drive a culture of automation and reduce manual activities.
  • Drive initiatives with partner teams to improve infrastructure reliability.

ServiceNow provides an AI platform for business reinvention, used by 85% of the Fortune 500. The company fosters an AI-native culture with a focus on innovation and collaboration.

Global

  • Be on an on-call rotation responding to production incidents and support service engineers.
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
  • Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.

Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.

US

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

$125,000–$135,000/yr

  • Manage patching, updates, and lifecycle for Azure VMs, support on-premises infrastructure transition, and maintain operational documentation.
  • Act as escalation point for complex infrastructure issues, troubleshooting cloud, server, networking, identity, and application problems.
  • Develop PowerShell scripts, automation solutions, and leverage AI-assisted tools to improve operational efficiency.

AIP Publishing is a leading publisher of the physical sciences, accelerating scientific discovery and reimagining scholarly publishing. It is a wholly owned not-for-profit subsidiary of the American Institute of Physics, fostering a collaborative and creative atmosphere that maximizes individual contributions.

Ireland

  • Investigate and resolve customer technical issues across cloud security posture management, vulnerability scanning, threat detection, and container/Kubernetes security.
  • Troubleshoot cloud connector and integration failures across AWS, Azure, GCP, OCI, and SaaS platforms.
  • Design and implement automation leveraging AI agents and tooling to improve triage accuracy and resolution efficiency.

Wiz is a cloud security platform that enables teams to secure cloud and AI applications by connecting code, cloud, and runtime into a single shared context. As one of the fastest-growing startups, powered by Google, the company is trusted by over 65% of the Fortune 100 and scans over 230 billion files daily.

$64,021–$92,683/yr
Canada

  • Extend the self-service datastore platform with provisioning automation, guardrails, and paved paths for product engineering teams.
  • Ship observability, alerting, and backup/disaster recovery as built-in defaults for every datastore.
  • Convert recurring pull-in work into platform features or AI tooling that other teams can use directly.

Greenhouse provides a hiring software platform designed to make hiring work for everyone. They have an award-winning culture recognized by Fortune and Inc., and foster inclusivity, transparency, and accountability among their teams.

US

  • Own escalations end-to-end, reproducing, isolating, and communicating progress to stakeholders.
  • Perform deep technical analysis using SQL forensics, logs, and telemetry to identify root causes.
  • Document repro kits and internal RCA notes while building the Knowledge-Centered Support library.

HungerRush is a leading provider of integrated restaurant technology solutions, including the cloud POS system HungerRush 360. The company values team, respect, accountability, customer, and speed, and fosters a collaborative culture.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

Global 6w PTO

  • Operate as the NOC’s first point of escalation for any issues raised within or out of shift that need further assistance or feedback.
  • Ensure NOC shift operations are executed properly and according to pre-defined SLAs.
  • Manage NOC’s shift schedule including PTO requests and act as focal point for administrative issues.

LivePerson is a global leader in enterprise conversations, providing a Conversational Cloud platform for brands like HSBC and Chipotle. The company powers nearly a billion interactions monthly and fosters an inclusive workplace culture that encourages collaboration and innovation.

US

  • Provide solutions to customers to make them successful using our products.
  • Troubleshoot customer environments and engage in active triaging with customers.
  • Participate in on-call rotation for weekend coverage.

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life with its unified DataOps platform, Astro, powered by Apache Airflow. Trusted by more than 800 enterprises, the company fosters a diverse and inclusive culture as an equal opportunity employer.

North America

  • Serve as the primary technical advisor for a portfolio of strategic customers.
  • Lead technical reviews to identify risks, optimization opportunities, and improvement plans.
  • Investigate complex technical issues, perform root cause analysis, and drive resolution.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains and supporting the people who manage them. With over 2000 employees worldwide and a best-in-class HQ in Ottawa, Canada, the company fosters a culture of innovation and has won multiple Top Employer awards.