Source Job

Global

  • Define and execute the long-term vision for infrastructure design, building resilient distributed systems for AI workloads.
  • Mentor junior engineers to foster a culture of high standards, code quality, and continuous learning.
  • Diagnose complex technical bottlenecks and break down challenges into scalable, actionable requirements.

Software Development Distributed Systems Cloud Infrastructure Kubernetes

20 jobs similar to Senior Staff Engineer

Jobs ranked by similarity.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

US 4w PTO 14w maternity 14w paternity

  • Lead a team of product engineers to build the core developer experience of the Render platform.
  • Drive execution on key product initiatives, balancing speed with high reliability and quality.
  • Partner with Product and Design to translate customer needs into technical solutions.

Render builds a developer-first cloud platform for AI-native applications, eliminating the complexity of hyperscalers. With over 6 million developers and $257M in funding, the team values craft, velocity, and user experience.

US 20w maternity 12w paternity

  • Own end-to-end technical execution for strategic customer and partner engagements, including discovery, infrastructure design, implementation, and production deployment.
  • Design and build cloud infrastructure supporting advanced AI workloads, including simulation, training, evaluation, inference, and large-scale batch processing.
  • Improve platform reliability, security, performance, and cost efficiency by debugging issues across application, network, storage, compute, and orchestration layers.

The partner company is building the infrastructure foundation for next-generation AI applications and physical AI workloads. The engineering team is pioneering and values ownership, technical excellence, and solving challenging engineering problems at scale.

US

  • Architect large-scale infrastructure systems for end-to-end software lifecycle platforms.
  • Design unified CI/CD and deployment platforms for build, test, deploy, and runtime operations.
  • Build observability, monitoring, and cost visibility systems for production infrastructure.

LinkedIn is the world's largest professional network, creating economic opportunity for every member of the global workforce. As a large company, LinkedIn fosters a culture of trust, care, inclusion, and fun, investing in employee growth.

$180,000–$200,000/yr
Global

  • Lead customers in designing and optimizing GPU-based solutions on Vultr's platform.
  • Collaborate with cross-functional teams to bring AI, ML, and GPU workloads into production.
  • Educate customers on the value of Vultr's cloud infrastructure and expand their possibilities.

Vultr provides high-performance cloud infrastructure solutions globally, making them easy to use, affordable, and locally accessible. It is a privately-held company with over a decade of self-funding, hundreds of thousands of customers across 185 countries, and a culture that emphasizes comprehensive benefits and employee growth.

$150,000–$175,000/yr
US

  • Design and manage GCP project structure, networking, and core infrastructure.
  • Own Kubernetes cluster infrastructure and define infrastructure-as-code standards.
  • Improve system architecture for scalability, resilience, and performance.

UJET provides an AI-powered contact center platform that redefines customer experience with cloud-native architecture and mobile-first approach. The company is a growing tech firm with a collaborative culture focused on innovation and security.

US Canada Unlimited PTO

  • Lead a unified platform strategy defining and executing a cohesive roadmap aligning Production Engineering with Tenant Experience.
  • Drive operational excellence and reliability, maintaining deep accountability for GitLab’s production outcomes and incident management.
  • Scale engineering leadership by managing a senior management team, coaching on performance, hiring, and team structure.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab is where careers accelerate and innovation flourishes.

US

  • Design, build, and operate scalable infrastructure platforms for large-scale AI model training and inference.
  • Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
  • Develop software solutions and automation tools using Python and systems programming languages like Go or C++.

Our partner builds and operates foundational technology powering advanced AI training and inference workloads at scale. They offer a collaborative culture focused on innovation, engineering excellence, and continuous learning.

$175,350–$250,500/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead a team of senior engineers building and operating Docker's secure build infrastructure.
  • Own delivery and turn an ambitious roadmap into a concrete plan for predictable execution.
  • Stay deeply engaged in technical work, from design discussions to code review and hands-on contributions.

Docker builds tools for developers to build, share, and run applications, trusted by over 20 million users. The company is a globally distributed, remote-first team with offices in Seattle and Paris.

$175,000–$229,000/yr
US

  • Lead end-to-end architecture and implementation of high-performance detection tools.
  • Mentor junior engineers through design reviews and promote maintainable code standards.
  • Collaborate with product managers and scrum masters to deliver high-quality releases on schedule.

Corelight defends the world's most sensitive networks by transforming digital footprints into actionable insights, using open-source innovations and agentic AI. It is a fast-growing network detection and response platform with a collaborative, inclusive, distributed culture.

$124,000–$329,200/yr
US

  • Design and scale highly available backend services and APIs supporting AI-powered developer tools.
  • Develop distributed systems optimizing reliability, latency, cost, and performance at global scale.
  • Provide technical leadership through mentorship, code reviews, and collaboration across engineering teams.

Our partner is building next-generation AI-assisted software development experiences. This remote-first role contributes to a globally impactful AI platform with a collaborative culture focused on innovation, inclusion, and technical excellence.

US Unlimited PTO

  • Architect and improve cloud foundations on Google Cloud Platform to support scalable, secure, and well-governed workloads.
  • Design and build platform capabilities across GCP, Kubernetes, CI/CD, GitOps, and developer tooling.
  • Mentor engineers and raise the technical bar through code review, architecture guidance, and direct implementation.

Wpromote is a digital marketing agency focused on performance marketing and technology. The company fosters a diverse, inclusive culture with a remote-friendly environment and office hubs in Los Angeles, Chicago, and New York.

Global Unlimited PTO

  • Lead a high-impact infrastructure team, evolving internal platforms and CI/CD systems to support large-scale engineering operations.
  • Drive automation initiatives and AI-driven practices to reduce operational complexity and improve developer experience.
  • Define and execute strategies for scalable infrastructure, cloud environments, and platform engineering.

The partner company is a technology organization focused on building infrastructure platforms that enable engineering teams to deliver software faster. It is a remote-first company with a collaborative culture and a focus on innovation and scalability.

$180,000–$215,000/yr
US Unlimited PTO

  • Lead the direction for detecting and responding to AI threats using creative solutions.
  • Collaborate cross-functionally with engineering, IT, legal, and product teams.
  • Manage globally distributed security teams with a focus on data and execution.

Gainsight is the AI-powered retention engine behind the world's most customer-centric companies. With over 2,000 customers and a culture focused on collaboration and human-first values, we offer stability and support for our team.

$180,000–$300,000/yr

  • Build and scale high-performing engineering teams, hiring and developing top technical talent.
  • Define and drive technical strategy, execution, and long-term roadmap alignment with business objectives.
  • Oversee architecture, infrastructure, and best practices for scalable, reliable systems.

SignalFire partners with top early-stage startups shaping the future of technology. They have a portfolio of over 200 companies across AI, cybersecurity, healthtech, fintech, developer tools, and enterprise SaaS.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

Global

  • Apply cloud infrastructure and platform engineering expertise to train next-generation AI systems.
  • Create realistic cloud infrastructure tasks involving distributed systems, networking, security, scalability, and reliability.
  • Build reproducible, containerized environments with valid reference solutions and intentionally defective variants.

Our client is a rapidly growing, venture-backed AI company helping shape the next generation of intelligent systems. Backed by more than $40 million in funding and supported by a rapidly expanding international network of experts, the company is building critical human intelligence infrastructure for the AI economy.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.