Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
Build and maintain observability systems and automate operational toil using AI tools.
Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
Monitor production systems, participate in incident response, and champion DevOps best practices.
First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.
ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.
Consolidate Terraform and establish conventions for state management, modules, and CI checks.
Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.
Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.
Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.
Contribute to platform and harness engineering, including CI/CD and developer tooling.
Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
Participate in on-call rotation and incident management to ensure platform uptime.
Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
Build and maintain automation for lifecycle management of Kubernetes clusters.
Identify and root-fix reliability bottlenecks before they become incidents.
Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
Balance 60% hands-on technical work with 40% people leadership and career growth.
Drive a maturing reliability practice including SLOs, incident response, and on-call.
Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
Lead incident response and implement observability to improve availability and performance.
Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.
Design, build, and operate high-scale observability pipelines for logs, metrics, traces, and exceptions.
Lead cross-functional initiatives to resolve scaling bottlenecks and evolve production infrastructure safely.
Partner with engineering teams to improve observability tools and provide technical leadership across teams.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It emphasizes learning, professional growth, and an inclusive work environment.