Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.
Acts as the strategic bridge between Cloud Operations, Product Management, Engineering, and other teams to drive service quality and operational excellence.
Drives large-scale transformation programs, promotes operational best practices, and ensures lessons learned translate into portfolio-wide improvements.
Champions automation, observability, and reliability standards while influencing engineering practices and product roadmaps.
Unit4 is a cloud company redefining ERP for mid-market people-centric organizations with over 40 years of heritage. They are a people-first community focused on trust, accountability, and growth, with a global team and a commitment to sustainability and inclusion.
Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.
Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.
Design, build, and operate production Kubernetes platforms, owning cluster architecture, networking, reliability, and security.
Troubleshoot complex infrastructure issues across Kubernetes, Linux, networking, and cloud environments, and improve observability and automation.
Take ownership of critical infrastructure initiatives, incident response, and mentorship while working autonomously in a remote-first environment.
Our partner is a technology company building and operating large-scale Kubernetes platforms and complex production infrastructure. It is a remote-first organization with a focus on autonomy, technical ownership, and collaboration across North America and Latin America.
Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
Build and maintain automation for lifecycle management of Kubernetes clusters.
Identify and root-fix reliability bottlenecks before they become incidents.
Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.
Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
Lead incident response and partner with product engineers to design resilient systems for enterprise customers.
Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.
Own and evolve the product roadmap for Observability and Compute/Network teams, from metrics and alerting to Kubernetes and capacity planning.
Partner with engineering and security leaders to prioritize reliability, cost, and developer experience.
Drive self-service adoption and track outcomes like SLO attainment, MTTD/MTTR, and infrastructure cost.
Addepar is a global data and AI platform that empowers investment professionals to turn complex financial information into actionable intelligence. More than 1,500 firms in 60 countries use Addepar to manage nearly $10 trillion in assets, with an inclusive, ownership-driven culture.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Lead a two-month EKS modernization discovery for a high-scale consumer mobile platform and deliver a prioritized roadmap.
Drive execution across blast radius reduction, automated upgrades, compute right-sizing, and Graviton migration.
Act as Pod Leader in the U.S.-Based Virtual Operating Center and be the primary technical contact for customer infrastructure leadership.
EverOps is a premier Embedded Service Provider partnering with customer engineering teams on mission-critical infrastructure and cloud challenges. The company has been fully remote since day one and hires senior engineers who drive meaningful outcomes.
Own infrastructure end to end, including Kubernetes clusters, Docker builds, and DigitalOcean services.
Move production deploys toward safer, more frequent releases and manage capacity, autoscaling, and load balancing.
Operate PostgreSQL, Redis, and NATS under real traffic, keep Cloudflare tight, and document systems for the whole team.
Colonist is building the biggest digital board game platform on the internet, where players have spent over 3,000 years playing 60 million games on their Catan alternative. The company is a fully remote, asynchronous team spread across multiple continents, focused on crafting polished digital board games.
Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
Own production infrastructure, containerized applications, deployment workflows, and monitoring.
Collaborate with development teams to streamline CI/CD and drive high availability.
Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.
Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).
DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Architect and maintain cloud infrastructure for F5's AI security platform, ensuring scalability and reliability.
Design CI/CD pipelines, manage Kubernetes workloads, and implement serverless technologies.
Build comprehensive monitoring, security, and compliance practices across the B2B SaaS solution.
F5 is a global leader in application delivery and security, helping organizations create, secure, and run applications. With over 6,400 employees and 553 patents, F5 serves more than 23,000 customers across 170 countries and fosters a human-first, inclusive culture.
Own end-to-end customer onboardings, from sales handoff to operational readiness.
Build and maintain success plans that drive adoption, renewal, and expansion.
Engage with platform engineers and executives to translate platform adoption into measurable business outcomes.
vCluster Labs is the leading platform for AI infrastructure, providing a hyperscaler-like experience on GPU infrastructure for AI cloud providers and enterprises. It is a venture-backed startup with over 40 engineers and a remote-first culture, trusted by more than 50 AI clouds and Fortune 500 companies.
Own and evolve production infrastructure, leading the migration from Docker Swarm to Kubernetes on premises.
Drive observability, enforce IaC practices, and ensure CI/CD reliability across ~50 services.
Participate in on-call rotation, resolve incidents, and build platform tooling to reduce infrastructure toil.
Webshare is an enterprise-grade proxy platform providing access to over 80 million global IPs across 195 countries. With 99.97% uptime, it serves tens of thousands of businesses, and the team emphasizes mentorship, knowledge-sharing, and team events.
Design and operate scalable AWS infrastructure with containerization and orchestration tools.
Implement monitoring, logging, and infrastructure as code using Terraform.
Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.
Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.
Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.
Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.