Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.
Build and maintain scalable cloud infrastructure for high availability.
Enhance observability and monitoring frameworks for accurate alerts.
Support on-call rotations and incident response with post-mortems.
GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.
Consolidate Terraform and establish conventions for state management, modules, and CI checks.
Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.
Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.
Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.
ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).
DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.
Lead the architecture and implementation of complex cloud solutions across AWS and GCP.
Drive cloud automation and optimization initiatives to improve scalability and reliability.
Provide technical leadership and mentorship to engineers while collaborating with global teams.
The company focuses on cloud infrastructure and platform engineering. They operate with global teams and emphasize automation, security, and reliability.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Own best practices for managing production infrastructure, including provisioning, scaling, configuration, capacity planning, and monitoring.
Build and maintain Kubernetes infrastructure at scale alongside Terraform-provisioned cloud resources.
Write custom automation and tooling in Go to reduce manual work and eliminate operational risk.
Turnkey is building the infrastructure for the autonomous economy, providing programmable guardrails that enable organizations to operate with autonomy and control. Founded by the team behind Coinbase Custody, it is a deeply technical, low-ego, high-agency team of experts in cryptography, security, and systems.
Design and build scalable, reliable cloud infrastructure on GCP and AWS.
Manage Kubernetes environments and infrastructure as code with Terraform.
Drive CI/CD automation, platform reliability, and developer self-service.
The company builds and operates scalable cloud infrastructure and internal developer platforms. It is a globally distributed, fully remote engineering team with a collaborative and inclusive culture.
Architect and automate scalable cloud environments across AWS and Azure using Terraform, Ansible, Helm, and CDK.
Serve as Linux subject matter expert, managing system builds, core services, and performance from kernel up.
Lead CI/CD pipelines, observability, security, and incident response to ensure platform reliability.
Fueled is a leading digital strategy, design, and engineering agency. The 300+ person team has designed and built hundreds of digital products for major brands like Google, Apple, and The New York Times, and thrives in a culture that values flexibility, creativity, and cutting-edge technology.
Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.
Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.
Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.
Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.
Operate, scale, and troubleshoot Bitsight's SaaS cloud infrastructure with focus on reliability, efficiency, and security.
Tackle complex system-level designs and proactively anticipate performance and scalability issues.
Pioneer self-optimizing infrastructure systems using AI, ensuring manual and staging validation before production deployment.
Bitsight is a cyber risk management leader transforming how companies manage exposure, performance, and risk. Over 3,500 customers and 600 teammates work across Boston, Raleigh, New York, Lisbon, Singapore, and remote locations.