Architect, deploy, and manage highly available, fault-tolerant cloud infrastructure across Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
Maintain and scale declarative infrastructure using Terraform across a multi-hundred-file estate, enforcing GitOps workflows with Atlantis.
Build, maintain, and optimize robust automated pipelines for continuous integration and delivery using GitHub Actions, Jenkins, and ArgoCD.
Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is a scrappy, nimble organization dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.
Own cloud infrastructure and Kubernetes environment, keeping it reliable, secure, and cost efficient.
Lead the team in using AI-assisted engineering to design, build, and operate platform infrastructure.
Manage and mentor a team of engineers while staying hands-on to contribute directly to the work.
Doma Technology provides solutions for lenders, real estate professionals, title agents, and homeowners that make closings simpler and more efficient. The company values an entrepreneurial, people-first culture with a focus on diversity, equity, and inclusion.
Own daily IT and platform operations, resolving access requests, deployments, and infrastructure tasks.
Manage cloud infrastructure on GCP and Cloudflare, CI/CD pipelines, and monitoring.
Collaborate with DevOps & Security lead to harden systems and scale the platform.
Centrifuge is building open infrastructure for real-world assets on blockchain, partnering with major financial institutions. We are a well-funded, small, high-trust team backed by leading investors, with over $1.7B in TVL.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Own the technical strategy for multi-ecosystem scaling, defining architecture for onboarding new language ecosystems.
Drive end-to-end remediation automation, leading redesign of CVE workflows to close the loop from detection to verified release.
Set platform-wide technical direction spanning package index, build pipelines, and orchestration tooling to serve customers and ecosystem teams.
Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. They serve Fortune 500 enterprises and global industry leaders, and are venture-backed by leading investors, fostering a culture of customer obsession and intentional action.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.
Provide expert architectural guidance and day-two operational expertise for GCP infrastructure supporting a national aviation safety platform.
Partner with customer operations teams and Google technical advisors to strengthen platform reliability, automation, and DevSecOps practices.
Take proactive ownership of complex issues, from Tier 3 troubleshooting to operational readiness, ensuring high availability and performance.
540 is a forward-thinking consulting firm that partners with government agencies to deliver innovative technology solutions for mission-critical systems, including aviation safety platforms. The company fosters a culture of ownership, collaboration, and technical excellence, with a small-team environment where engineers drive real impact.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
Partner cross-functionally to align platform priorities with business goals and ensure system reliability.
PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.
Own infrastructure as code across development, staging, and production environments
Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
Manage cloud infrastructure, establish scalable engineering practices, and lead incident response
CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.
Design and evolve cloud infrastructure on GCP for scale and resilience.
Build internal tooling and automation that promote team autonomy and developer productivity.
Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.
The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.
Design and evolve platform architecture, balancing scalability, reliability, and security.
Build tooling, services, and automation that improve developer productivity and observability.
Own infrastructure as code, CI/CD pipelines, and event-driven systems across a cloud-native stack.
Pivotal Health is a technology platform that helps healthcare providers get paid fairly through AI-driven reimbursement workflows. The company is a collaborative, low-ego team on a mission to make healthcare reimbursement fairer for providers.
Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
Build the human-AI development loop, creating tooling and automation for coding agent workflows.
Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.
Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.
Lead and mentor a distributed engineering team across US and EU, fostering growth and collaboration.
Drive proactive ownership and engineering excellence for a cloud platform at exabyte scale.
Architect global infrastructure expansions across multi-cloud and compliance environments.
New Relic is an intelligent observability platform that helps companies gain insight into their complex systems and thrive in an AI-first world. As a global team of innovators, we foster a diverse, welcoming, and inclusive environment where employees can be their authentic selves.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Design and evolve cloud architecture on GCP, expressing it entirely as code with Terraform following GitOps principles.
Build and own CI/CD pipelines for IaC, including Policy-as-Code guardrails, drift detection, and progressive rollout.
Advance Platform-as-a-Product by building self-serve capabilities so engineers can provision what they need.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and more. With over 400 globally distributed team members and $400 million in funding from top-tier investors, Alpaca is committed to open-source contributions and fostering a vibrant community.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.