Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Build and maintain scalable cloud infrastructure for high availability.
Enhance observability and monitoring frameworks for accurate alerts.
Support on-call rotations and incident response with post-mortems.
GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.
Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.
Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Consolidate Terraform and establish conventions for state management, modules, and CI checks.
Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.
Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.
ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.
Contribute to platform and harness engineering, including CI/CD and developer tooling.
Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
Participate in on-call rotation and incident management to ensure platform uptime.
Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.
Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.
Branch empowers workers with financial freedom by helping companies accelerate payments and providing accessible, free financial services. It is a remote-first, award-winning FinTech with employees across the U.S., fostering a culture of transparency, accountability, and trust.
Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
Diagnose and troubleshoot complex distributed systems running at high request volume.
Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.
Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
Lead incident response and partner with product engineers to design resilient systems for enterprise customers.
Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.
Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
Implement observability solutions and enforce security best practices to ensure uptime and system performance.
Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.
Own the reliability posture of production services, including availability, latency, capacity, and performance.
Define and operate against SLIs and SLOs, using error budgets to drive engineering priorities.
Lead incident response, write post-mortems, and build automation to measurably improve service reliability.
Twilio is a cloud communications platform that delivers innovative solutions to hundreds of thousands of businesses and empowers millions of developers worldwide to create personalized customer experiences. The company is remote-first with a strong culture of connection, global inclusion, and a focus on solving problems and taking initiative.