Architect, deploy, and manage highly available, fault-tolerant cloud infrastructure across Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
Maintain and scale declarative infrastructure using Terraform across a multi-hundred-file estate, enforcing GitOps workflows with Atlantis.
Build, maintain, and optimize robust automated pipelines for continuous integration and delivery using GitHub Actions, Jenkins, and ArgoCD.
Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is a scrappy, nimble organization dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.
You will own end-to-end data protection: architecting security infrastructure, managing PCI/SOC2 programs, and setting the security roadmap.
You will integrate security scanning (SAST, DAST, SCA) into CI/CD pipelines and harden containerized and cloud-native environments.
You will configure IAM and SSO platforms, manage EDR/DLP/SIEM tooling, and run security awareness training.
Campminder builds software for summer camps, enabling meaningful experiences for kids. With over 100 employees, they are stable, profitable, and have a values-led culture committed to work/life balance.
Own the stability and deployment of Supabase's PostgreSQL products, acting as a bridge between product and infrastructure teams.
Manage PostgreSQL lifecycles, package software using Nix, and optimize CI/CD tooling for self-service releases.
Resolve production issues by proactively identifying and fixing customer deployment problems while maintaining best practices and tests.
Supabase is the Postgres development platform, providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. The company has around 400 team members across 60+ countries, is open-source-first and fully remote, with a culture of async work and building developer tools at scale.
Design and implement production systems and data center networks with a focus on automation and reliability.
Develop and maintain infrastructure automation using tools like Terraform and Ansible, and script in Python and Bash.
Troubleshoot and resolve complex issues across the full technology stack, including Linux, Windows, and AWS environments.
Our client is a global technology company transforming marketing decisions through a platform that enables cross-channel programmatic media campaigns. The company operates across multiple continents and partners with over 150 major platforms including Facebook, Instagram, and Twitter.
Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
Automate operations and engineering tasks to improve productivity and reliability.
Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.
Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.
Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.
Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information.
Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
Support development teams by consulting on infrastructure needs and improving observability and incident response.
Provide operational support and maintain platform reliability through structured debugging and on-call rotations.
PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.
Ensure reliability, scalability, and performance of cloud-based systems using Kubernetes and observability tools.
Define and monitor reliability metrics (SLIs, SLOs, MTTR) to continuously improve operational performance.
Automate operational tasks and implement Infrastructure as Code to reduce manual work and enhance efficiency.
Our partner is a technology company focused on building and maintaining reliable, scalable digital environments. They promote a culture of continuous improvement, collaboration, and proactive engineering.
Design, implement, and manage scalable cloud infrastructure solutions using Azure and AWS.
Develop and maintain infrastructure as code with Terraform and manage Kubernetes containerized workloads.
Build CI/CD pipelines and apply site reliability engineering practices to improve system reliability and performance.
They specialize in cloud infrastructure and DevOps solutions for enterprise applications. The company operates with a distributed team and offers a flexible, fast-paced work environment.