Source Job

  • Own end-to-end infrastructure deployment programs for new capacity and site expansions, including hardware dependencies and commissioning gate frameworks.
  • Deliver crisp, data-driven executive updates and govern cross-organizational dependencies without escalation.
  • Coach junior TPMs and drive AI tool integration to improve program tracking and risk detection.

Technical Program Management AI Tools Networking

20 jobs similar to Staff Technical Program Manager, Deployments

Jobs ranked by similarity.

$125,000–$175,000/yr

  • Scale spatial AI deployments from pilot to hundreds of stores per retail partner.
  • Own the end-to-end technical plan spanning CV, hardware, and software.
  • Drive metrics-driven rigor and translate technical trade-offs for customers.

Augmodo builds spatial AI systems for retail, managing technical rollouts that combine hardware, firmware, and computer vision pipelines. They foster a pragmatic, inclusive culture with a focus on measurable execution and customer success.

$140,000–$165,000/yr
Global Unlimited PTO

  • Partner with Data Center Management and Host teams to run capacity review cadences that feed Supply's input into the product roadmap.
  • Own the relationship and escalation path with host and infrastructure partners, resolving capacity, performance, or contractual issues quickly.
  • Translate capacity and utilization data into concrete tradeoffs on cost and reliability, and drive project plans across Engineering, Product, and Supply stakeholders.

Runpod is the AI Developer Cloud, serving over one million developers in building and scaling AI models. We are a small, remote-first team that values ownership, speed, and impact.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

US Unlimited PTO

  • Manage provider relationships, onboarding, and incident response for AI infrastructure.
  • Act as incident commander during provider-side failures, coordinating SREs, providers, and customers.
  • Build playbooks and standards to scale provider onboarding and operations.

Andromeda Cluster gives early-stage startups access to scaled AI infrastructure. We are a growing team building the liquidity layer for global AI compute.

Canada

  • Lead cross-functional development of large-scale infrastructure programs from concept to delivery.
  • Align stakeholders across engineering, product, and finance to manage risks and drive program success.
  • Optimize engineering processes and drive cost efficiency through automated monitoring and reduction initiatives.

Lime is a global shared micromobility company on a mission to make transportation shared, affordable, and carbon-free. It has powered over one billion rides in nearly 30 countries and is a Time Magazine 100 Most Influential Company, fostering a culture of strategic thinking and operational excellence.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

US

  • Review, design, and vet floor plans for high-density AI data centers, ensuring power, cooling, and network layouts meet requirements.
  • Manage the full deployment lifecycle from pre-approval site vetting through design, implementation, and operational handoff.
  • Coordinate across internal teams and external vendors to align on timelines, milestones, and risk mitigation.

Megaport is a global leader in Network as a Service (NaaS) that connects businesses to cloud, data centers, and each other. With over 600 employees across Asia-Pacific, Europe, and the Americas, the company fosters a collaborative, supportive, and fun culture.

Global Unlimited PTO

  • Own the technical evaluation end-to-end, from discovery to POC, ensuring evaluations are scoped and tied to customer ROI.
  • Take customers from signature to first successful production training run and serve as the technical owner post-launch.
  • Build the SA function by creating demo environments, benchmarking harnesses, and reference architectures.

Andromeda Cluster gives early-stage startups access to scaled AI infrastructure, partnering with leading AI labs and cloud providers. It is a high-growth company building an inclusive environment for all employees.

Global 6w PTO 26w maternity 26w paternity

  • Manage the program portfolio covering inference, efficiency, serving, and endpoints to scale Cohere's infrastructure.
  • Lead cross-functional coordination with Modeling and customer-facing teams for end-to-end execution.
  • Identify pain points, establish processes, and improve engineering best practices.

Cohere is the leading security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are a global team of researchers, engineers, and designers passionate about our craft, with offices across North America and Europe.

Europe

  • Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
  • Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
  • Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. It is a fast-growing scale-up with a global team across the USA, Australia, Central Europe, Malaysia, Singapore and Japan.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

US Unlimited PTO

  • Leads the most complex product development programs for data center infrastructure, driving products from concept through scaled production.
  • Defines end-to-end manufacturing strategy, including make-vs-buy decisions, partner selection, and capacity planning.
  • Drives cross-functional leadership, quality improvements, and operational rigor across internal and external stakeholders.

Tract Capital adopts a unique approach to digital infrastructure investment, leveraging decades of experience to nurture leading-edge digital infrastructure enterprises. Their team of specialized experts acts as a strategic partner and catalyst for innovation within the sector.

Canada

  • Lead, coach, and develop a team of Deployment Specialists and Managers, providing clear direction and performance management.
  • Own delivery standards across numerous teams, including tooling, documentation, and quality benchmarks.
  • Oversee scheduling, staffing, and capacity planning to ensure delivery commitments are met without risk to team sustainability.

Versaterm is a PE-backed, high-growth SaaS company building a fully integrated platform across the public safety ecosystem. They have grown through strategic acquisitions and deep product investment, attracting people who want to be pushed by great people and build impactful solutions.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

$180,000–$260,000/yr
US

  • Drive strategic programs through the full Concept, Commit, and Execution lifecycle across R&D, Engineering, Product Management, and Sales.
  • Keep programs aligned across functions, connecting technical delivery to commercial outcomes, and ensuring each team knows its commitments.
  • Present program status, risks, and decision points to executive leadership with clarity and candor, framing options rather than surprises.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many leading enterprises globally and is committed to open standards and technical excellence, fostering a culture of collaboration and innovation.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

$150,000–$200,000/yr
US Unlimited PTO

  • Lead end-to-end execution of new product development programs from concept through production.
  • Partner with Engineering, Manufacturing, Supply Chain, Quality, Construction, Commissioning, and Operations to align priorities.
  • Proactively identify program risks, resource constraints, and execution gaps while driving mitigation plans.

Tract Capital is a digital infrastructure investment firm that leverages decades of experience to nurture and advance leading-edge enterprises. The company has a team of specialized experts united by a singular purpose: to support the growth of digital infrastructure.

Germany

  • Coordinate complex engineering initiatives supporting the delivery and scalability of advanced AI infrastructure and platforms.
  • Drive cross-functional projects from planning through execution, managing dependencies, risks, and stakeholder alignment.
  • Continuously improve project management methodologies and engineering workflows to enhance predictability and efficiency.

The company focuses on advanced AI infrastructure and platforms, enabling cutting-edge technologies. It fosters a collaborative, international, and fast-paced remote culture where technical excellence and ownership are highly valued.

$200,000–$350,000/yr
US

  • Work directly with customers to understand technical requirements and build production solutions.
  • Develop integrations, APIs, and internal tooling to accelerate platform adoption.
  • Own projects from technical discovery through deployment and iteration.

A fast-growing technology company building advanced AI-powered infrastructure and products for enterprise customers. The company operates in a fast-paced, ambiguous environment and values high ownership and customer focus.