Design and execute performance benchmarks for AI training and inference workloads.
Profile and characterize GPU workloads to identify bottlenecks and optimization opportunities.
Systematically tune workload parameters to maximize throughput and establish performance baselines across GPU platforms.
Vultr provides high-performance, affordable cloud infrastructure for enterprises and AI innovators with 33 global data centers. It is the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, with a culture focused on growth and innovation.
Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.
Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
Create and maintain technical documentation and deliver training sessions, webinars, and demos.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.
Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.
Own end-to-end technical execution for strategic customer and partner engagements, including discovery, infrastructure design, implementation, and production deployment.
Design and build cloud infrastructure supporting advanced AI workloads, including simulation, training, evaluation, inference, and large-scale batch processing.
Improve platform reliability, security, performance, and cost efficiency by debugging issues across application, network, storage, compute, and orchestration layers.
The partner company is building the infrastructure foundation for next-generation AI applications and physical AI workloads. The engineering team is pioneering and values ownership, technical excellence, and solving challenging engineering problems at scale.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Lead the design and implementation of production-grade cloud and AI/ML solutions on AWS, ensuring scalability, security, and cost efficiency.
Act as a trusted technical advisor for customers, optimizing reliability, performance, security, and FinOps outcomes.
Deliver structured engagements such as architecture reviews, cloud optimization assessments, and GenAI workshops.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share shortlists with employers, operating with a global, distributed team.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Embed with enterprise customers to architect and deploy production-grade AI solutions on Google Cloud.
Design and implement RAG pipelines, multi-agent systems, and agentic wrappers using Vertex AI and Gemini.
Co-build with customer engineering teams and drive handoff for long-term ownership.
Valtech is the experience innovation company that helps brands unlock new value in a digital world. They foster a workplace culture of creativity, diversity, and autonomy with a global team.
Own and grow strategic customer relationships, driving AI infrastructure revenue as a trusted advisor.
Collaborate cross-functionally with Product, Solutions Engineering, and Customer Success to align on product capabilities and long-term success.
Maintain accurate sales data in CRM tools and leverage the AI partner ecosystem to enhance value propositions.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage, with 33 global data centers and customers in 185 countries. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr fosters a culture of innovation and self-funding.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Build customer-centric demo content and deliver hands-on, guided training sessions at customer sites globally.
Transition from engineering-centric demos to relatable, use-case-driven flows, and maintain a library of enablement artifacts.
Act as a credible technical voice, demonstrating GPU workload provisioning and cluster management on the platform.
Armada is a hyperscaler for the edge, delivering modular AI infrastructure from first deployment to AI factory. With nearly $500 million in funding and backing from top investors including Founders Fund, Lux, BlackRock, and Microsoft, the company is building the infrastructure layer for sovereign and edge AI.
Lead the architecture strategy for generative and agentic AI platforms, transforming complex business challenges into scalable solutions.
Collaborate with engineering teams, executives, and enterprise clients to define secure, production-ready architectures.
Influence AI adoption strategies, establish best practices, and shape the future of AI-powered solutions.
The company designs and delivers enterprise-grade AI solutions built on advanced cloud technologies. It offers a remote-first culture with a focus on innovation and collaboration among talented engineers.