Design, develop, and maintain CI/CD pipelines using Golang, Python, and Terraform to automate engineering processes.
Build and operate cloud-native platforms with Kubernetes and Docker, ensuring high reliability and scalability.
Collaborate with teams to improve automation, observability, security, and production support across services.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses an automated system to review applications and share top-fitting candidates with employers, focusing on efficiency and fairness in the hiring process.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design and build infrastructure primitives that define how CI/CD, build systems, and developer environments scale across the engineering org.
Build and operate the Kubernetes-based control plane behind CI/CD, including GitHub Actions runners, GitOps workflows, and ephemeral environments.
Develop core infrastructure components like Kubernetes Operators and scaling automation that product teams use directly, reducing bespoke per-team tooling.
Chainlink is the industry-standard oracle platform that brings capital markets onchain and powers the majority of decentralized finance. The company has enabled tens of trillions in transaction value and is adopted by major financial institutions and top protocols.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.
Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.
Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.
SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.
LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.
You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
You will respond to support escalations and troubleshoot complex technical problems.
You will develop software, tools, and scripts to automate deployment and monitoring of production systems.
We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.
Design and deliver solutions for cloud hosted production infrastructure.
Shape how mission-critical enterprise software solutions are developed and deployed using optimized CI/CD pipelines.
Design, build and support infrastructure and security technologies within the cloud.
Ping Identity provides an intelligent cloud identity platform that enables secure and seamless digital experiences. They serve over half of the Fortune 100 companies and have a global team that values diversity and individuality.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Optimize new and existing systems by increasing reliability, performance, and scalability.
Automate routine operational tasks to reduce toil and improve efficiency.
Ensure infrastructure security compliance and implement least-privilege access controls.
Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.
Ensure reliability, performance, and scalability of Backcountry's multi-cloud platform.
Drive incident resolution, postmortems, and automation to reduce operational toil.
Leverage AI-assisted engineering tools and collaborate with teams to build and maintain observability and SLI/SLO instrumentation.
Backcountry is an online retailer of outdoor gear and apparel, rooted in adventure and the outdoor lifestyle. The company fosters a culture of recognition, wellbeing, and connection, with a lean, fast-paced engineering team.
Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.
Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.
Ensure availability, performance, scalability, and resilience of production services in AWS.
Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
Collaborate with development, architecture, security, and product teams to promote reliability best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.