Drive execution of multiple infrastructure programs including network builds, storage initiatives, and datacenter expansions.
Own program planning, milestone definition, and dependency management across parallel workstreams.
Partner with engineering, operations, and leadership to ensure predictable execution in a fast-moving environment.
Lightning AI builds an end-to-end platform for developing, training, and deploying AI systems, based on PyTorch Lightning. They serve solo researchers to large enterprises, have offices in NYC, SF, Seattle, and London, and are backed by major VC firms.
Lead cross-functional technical programs spanning web platforms, data tooling, and AI-enabled systems.
Manage delivery, dependencies, and risks across multiple parallel initiatives in fast-paced environments.
Drive planning, documentation, launch readiness, and continuous improvement of program execution.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies, focusing on efficiency and objectivity. The company operates remotely with a small team and values a results-driven culture.
Drive large-scale, cross-functional new product development programs for APEX Strategic Initiatives, translating business objectives into execution strategy.
Develop trusted relationships with stakeholders and functional leaders to ensure program delivery and accountability.
Operate as an AI-native program manager, using AI tools to synthesize intelligence, surface risks, and provide actionable insights.
ServiceNow is the AI control tower for business reinvention, enabling 85% of the Fortune 500 to work smarter, faster, and better. The company fosters an AI-native culture where technology and talent work together, and is building a large, collaborative team focused on innovation.
Lead the delivery of complex product development programs driving growth and operational excellence.
Champion AI-native ways of working, using AI tools to accelerate program delivery and surface insights.
Establish governance frameworks and communication strategies to keep stakeholders informed and engaged.
ServiceNow is an AI platform company that helps businesses automate workflows and drive reinvention. It serves 85% of the Fortune 500 and fosters an AI-native culture where technology and talent are unstoppable.
Lead end-to-end infrastructure programs, breaking down large initiatives into manageable deliverables and measurable outcomes.
Proactively manage cross-team dependencies, risks, and blockers while aligning Infrastructure, Security, Product, and Engineering teams.
Use delivery metrics and tooling data to identify bottlenecks, improve workflows, and measure the impact of platform initiatives.
Prima is an online motor insurance provider that uses data and tech to deliver a great experience at a great price. They have over 5 million drivers, a team of 350+ engineers, and are expanding to the UK and Spain, fostering a culture of curiosity, experimentation, and collaboration.
Manage program management for the Security organization, applying a single operating standard across all projects.
Own intake, triage, status reporting, capacity planning, and prioritization using AI-augmented systems.
Partner with security leadership, engineers, and the AI team to refine and drive AI-assisted program management.
JumpCloud is the AI-powered unified IT management platform designed to secure the modern workforce. The company is remote-first and fosters a culture of connection and innovation, with teams in 15+ countries.
Manage the program portfolio covering ML inference, efficiency, serving, and endpoints to scale infrastructure for a growing user base.
Lead cross-functional collaboration with modeling and customer-facing teams to coordinate execution across programs.
Establish processes to improve engineering best practices, incident management, and prioritization to meet company priorities.
Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. It is a global team of researchers, engineers, and designers with offices in Toronto, San Francisco, London, and more, fostering a culture of passion and continuous improvement.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Design and implement a program-tracking system across concurrent R&D programs with clear milestones and risk visibility.
Build a prioritization framework that makes trade-offs explicit when new customer-driven scope arrives.
Map cross-team dependencies across hardware, ML, and software to prevent blockers and drive deployment efficiency.
Gather AI develops a vision-powered platform using autonomous drones to digitize warehouse workflows, improving efficiency and safety. They are a growing robotics company with a fast-moving, technically deep engineering culture.
Own the department's end-to-end operating rhythm, including quarterly planning, OKR cycles, roadmap management, and executive reviews.
Develop strategic communications and executive-level materials, including presentations and planning documents.
Drive transformation from AI-assisted workflows to AI-native operating practices by identifying automation opportunities.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use a remote-first work environment and provide support for productive home offices.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Own observability end to end and define how we measure reliability.
Own CI/CD pipelines and make shipping fast and safe.
Footprint builds Percy, an AI agent that runs financial crime investigations end to end. The company is backed by QED, Index, and other investors, and its small, senior team ships fast and grew revenue 5x in the past year.
Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.
Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.
Own the Claude Corps program plan and maintain the master program plan across all workstreams.
Manage day-to-day relationships with external partners (Anthropic, host organizations) and drive cross-functional alignment.
Codify and systemize mature workstreams to enable scaling from first cohort to thousands of fellows.
CodePath is the largest educator of college computer science students in the country. They have trained over 40,000 students from 1,000+ universities and partner with Amazon, Google, and other leading tech companies, and are scaling a $150M AI workforce program with Anthropic.
Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.
Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.
Lead end-to-end capacity planning and forecasting initiatives across large-scale production and infrastructure environments.
Develop and standardize operational processes, documentation, and workflows to enhance scalability and efficiency.
Manage complex technical programs involving multiple engineering teams, ensuring successful delivery against milestones and business objectives.
The company is a partner organization that focuses on providing technical program management and capacity planning services for enterprise clients. It operates as a fully remote team with a collaborative and dynamic culture, emphasizing innovation and continuous improvement.
Evolving the AI knowledge platform with retrieval, indexing, and synthesis for organization-wide use.
Architecting and operating agentic infrastructure on AWS with cost guardrails and observability.
Partnering with product engineering to define the AI platform API surface and building reference agent implementations.
ShiftKey is a healthcare workforce marketplace that connects facilities with licensed professionals to fill shifts, addressing staffing shortages. The company fosters an inclusive and collaborative culture, valuing diverse perspectives.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.