Ensure availability, performance, scalability, and resilience of production services in AWS.
Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
Collaborate with development, architecture, security, and product teams to promote reliability best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.
Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.
Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.
Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.
We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.
Ensure system architecture meets technical requirements by collaborating with IT teams (Architecture, Security, Infrastructure).
Maintain and evolve the microservices environment on AWS with a focus on information security.
Implement DevOps practices, automation, and monitoring tools to ensure system reliability and scalability.
Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 25,200 employees across 32 countries, it fosters a people-centric, inclusive culture recognized as a World's Best Workplace.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).
Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.
Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.
Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Act as a technical reference for SRE, DevOps, and cloud infrastructure, analyzing cloud environments in GCP/AWS for improvements and cost optimization.
Implement FinOps strategies, manage CI/CD pipelines, and maintain infrastructure as code using Terraform and Kubernetes.
Provide consultative support and communicate technical recommendations to engineering and business stakeholders.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses technology to streamline the application process and promote fair evaluation.
Optimize new and existing systems by increasing reliability, performance, and scalability.
Automate routine operational tasks to reduce toil and improve efficiency.
Ensure infrastructure security compliance and implement least-privilege access controls.
Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.