Source Job

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

Site Reliability Engineering DevOps Cloud Infrastructure CI/CD Infrastructure As Code

20 jobs similar to Platform Site Reliability Engineer

Jobs ranked by similarity.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

$160,000–$185,000/yr
US Unlimited PTO

  • Actively identify, plan and implement developer tooling and automation.
  • Participate in on-call rotations and assist with diagnostics and troubleshooting of platform and infrastructure.
  • Set technical direction for the team's infrastructure decisions and define overall DevOps strategy.

Bluesight creates groundbreaking solutions that increase efficiency, safety and visibility for health systems, hospital pharmacy, and pharmaceutical manufacturers. They are a high-growth healthcare information technology company with over 3,000 customers and a startup culture.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

Europe Unlimited PTO

  • Design and build scalable, reliable cloud infrastructure on GCP and AWS.
  • Manage Kubernetes environments and infrastructure as code with Terraform.
  • Drive CI/CD automation, platform reliability, and developer self-service.

The company builds and operates scalable cloud infrastructure and internal developer platforms. It is a globally distributed, fully remote engineering team with a collaborative and inclusive culture.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

Brazil

  • Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
  • Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
  • Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

$150,000–$170,000/yr
US 3w PTO

  • Lead cloud security, compliance, and infrastructure automation across AWS, Kubernetes, and CI/CD pipelines.
  • Translate HIPAA and SOC 2 requirements into practical technical controls and support audits.
  • Collaborate with Engineering and Data teams to automate workflows and maintain reliable system monitoring.

Synapticure is the largest specialty care and life sciences company for neurodegenerative diseases like Alzheimer's, Parkinson's, and ALS, operating in all 50 states. Founded by patients and caregivers, it is a remote-first company with a team across the US, driven by a mission to transform care and research.

UK 4w PTO

  • Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
  • Own production infrastructure, containerized applications, deployment workflows, and monitoring.
  • Collaborate with development teams to streamline CI/CD and drive high availability.

Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.

Ireland

  • Own the reliability posture of production services, including availability, latency, capacity, and performance.
  • Define and operate against SLIs and SLOs, using error budgets to drive engineering priorities.
  • Lead incident response, write post-mortems, and build automation to measurably improve service reliability.

Twilio is a cloud communications platform that delivers innovative solutions to hundreds of thousands of businesses and empowers millions of developers worldwide to create personalized customer experiences. The company is remote-first with a strong culture of connection, global inclusion, and a focus on solving problems and taking initiative.

UK 5w PTO

  • Design, implement, and maintain cloud infrastructure on AWS.
  • Develop and maintain CI/CD pipelines and deployment architectures.
  • Automate infrastructure provisioning and ensure security best practices.

Nearform is an independent team of data & AI experts, engineers, and designers who build intelligent digital solutions and capability at pace. Our team of 500 experts in 20+ countries is trusted by leading enterprises including Lululemon, Puma, Sun Life, Starbucks, and Walmart.

$250,000–$280,000/yr
United States

  • Lead both Site Reliability Engineering and Corporate IT, shaping technology operations strategy for a global SaaS environment.
  • Oversee platform reliability, incident response, compliance, and automation to reduce manual effort and improve efficiency.
  • Partner with Security and Compliance teams to support FedRAMP authorization and NIST 800-53 compliance frameworks.

Our partner is a high-growth, global SaaS environment seeking a Vice President of Enterprise Technology. They offer a fully remote executive role with approximately 10% travel and a focus on operational excellence.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Global

  • Design, provision, and maintain secure and scalable cloud infrastructure on AWS using EKS and ECS.
  • Build and optimize CI/CD pipelines using GitLab CI and Azure DevOps.
  • Utilize Terraform and Ansible for infrastructure as code and configuration management.

Miratech is a global IT services and consulting company that brings together enterprise and start-up innovation. The company retains nearly 1000 full-time professionals with a culture of Relentless Performance.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

$150,000–$182,400/yr
US 4w PTO

  • Design, develop, and maintain scalable, secure software services with an emphasis on backend systems, microservices, and APIs. - Implement DevSecOps practices including CI/CD pipelines, Infrastructure as Code, and containerization. - Collaborate with cross-functional teams to improve security, reliability, and delivery of healthcare technology solutions.

Bellese is a mission-driven digital services company pioneering innovative technology solutions in civic healthcare. They foster a collaborative, remote-first culture with a focus on learning and making a meaningful impact on public health outcomes.

$165,000–$200,000/yr
US

  • Design, build, and maintain cloud infrastructure on GCP and AWS using Terraform, optimizing CI/CD pipelines for rapid deployments.
  • Implement comprehensive observability including monitoring, logging, alerting, and distributed tracing to ensure platform health.
  • Establish and enforce security best practices, support AI/ML infrastructure, and build developer experience tooling.

Re:Build operates an advanced, end-to-end manufacturing platform that partners with industrial companies to bring products from concept to full-scale production. The company is guided by The Re:Build Way principles and aims to revitalize America's manufacturing base, creating meaningful jobs across the country.