Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Develops and maintains scalable cloud infrastructure on Google Cloud Platform to support AI applications and robotic systems.
Designs and implements APIs and backend services that enable seamless integration between cloud platforms, web applications, and robotics.
Collaborates with machine learning and research teams to optimize data processing, system performance, and reliability.
This partner company is at the forefront of AI and autonomous technology, developing advanced systems that power the future of robotics and intelligent automation. The team is collaborative and innovative, focused on building scalable cloud infrastructure and AI applications.
Design and scale highly available backend services and APIs supporting AI-powered developer tools.
Develop distributed systems optimizing reliability, latency, cost, and performance at global scale.
Provide technical leadership through mentorship, code reviews, and collaboration across engineering teams.
Our partner is building next-generation AI-assisted software development experiences. This remote-first role contributes to a globally impactful AI platform with a collaborative culture focused on innovation, inclusion, and technical excellence.
Own end-to-end technical execution for strategic customer and partner engagements, including discovery, infrastructure design, implementation, and production deployment.
Design and build cloud infrastructure supporting advanced AI workloads, including simulation, training, evaluation, inference, and large-scale batch processing.
Improve platform reliability, security, performance, and cost efficiency by debugging issues across application, network, storage, compute, and orchestration layers.
The partner company is building the infrastructure foundation for next-generation AI applications and physical AI workloads. The engineering team is pioneering and values ownership, technical excellence, and solving challenging engineering problems at scale.
Partner with engineering teams to improve reliability, scalability, and operational health of production systems.
Investigate and resolve complex production incidents, designing sustainable long-term solutions.
Design, build, and maintain automation tools and infrastructure to enhance developer productivity.
The company builds and maintains highly reliable, scalable production systems supporting millions of users worldwide. It fosters a remote-first culture that values innovation, collaboration, and engineering excellence.
Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
Create and maintain technical documentation and deliver training sessions, webinars, and demos.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Design, develop, and enhance backend systems for enterprise AI governance.
Build scalable services and contribute to architectural decisions.
Collaborate with product and engineering teams to deliver impactful features.
The company builds a governance platform that helps organizations manage and scale responsible AI initiatives. Its size is not specified, but it fosters a collaborative, innovative culture focused on learning, continuous improvement, and diversity.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Design, build, and maintain scalable machine learning infrastructure on AWS, including training and deployment pipelines.
Develop and deploy ML models for recommendation systems, fraud detection, credit risk, and personalization use cases.
Implement monitoring, logging, and alerting systems to ensure model performance, stability, and reliability in production.
Our partner is a fast-growing, innovation-driven company where machine learning and AI systems directly power large-scale fintech and commerce experiences. They foster a highly dynamic environment with strong emphasis on experimentation, rapid iteration, and measurable business impact.
Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks.
Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement.
Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company offers innovative solutions with integrated intelligence on a single enterprise platform, connecting the pharmacy ecosystem.
Design, build, and maintain foundational AI platform components including model gateways, vector databases, and LLM evaluation frameworks.
Develop secure platform architectures with integrated authentication, authorization, and access control mechanisms.
Build and enhance observability solutions for AI workloads and partner with teams to integrate AI infrastructure with agent frameworks.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. It processes applications and shares shortlists with hiring companies, operating as an intermediary in the recruitment process.
Architect and evolve scalable backend systems using Python and modern technologies.
Build distributed, data-intensive systems with high performance and operational excellence.
Collaborate with product, data, and machine learning teams to integrate AI/ML capabilities.
This partner company builds large-scale, data-intensive systems for next-generation digital experiences. It fosters a collaborative, remote-first culture focused on innovation, performance, and long-term impact.
Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.
Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.
Lead customers in designing and optimizing GPU-based solutions on Vultr's platform.
Collaborate with cross-functional teams to bring AI, ML, and GPU workloads into production.
Educate customers on the value of Vultr's cloud infrastructure and expand their possibilities.
Vultr provides high-performance cloud infrastructure solutions globally, making them easy to use, affordable, and locally accessible. It is a privately-held company with over a decade of self-funding, hundreds of thousands of customers across 185 countries, and a culture that emphasizes comprehensive benefits and employee growth.
Learn and contribute to production infrastructure automation using Ansible and Terraform.
Assist in building and maintaining cloud-native services across VKE, VLB, and VCR.
Develop hands-on skills with Kubernetes and container runtimes through guided project work.
Vultr provides high-performance cloud infrastructure solutions including Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage for enterprises and AI innovators worldwide. The company is the world's largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries, and values innovation and employee growth.
Build agentic applications that solve real customer use cases using our platform capabilities.
Design agentic workflows, plugins, and tool integrations that deliver measurable customer value.
Translate ambiguous problems into intuitive AI experiences and rapidly iterate into production-ready solutions.
Legion builds secure, reliable AI systems for government and enterprise, born from a Department of Defense partnership. Trusted by leaders like Palantir, Nvidia, HPE, and Oracle, the company focuses on embedding intelligence into complex systems to optimize workflows.
Lead the architecture strategy for generative and agentic AI platforms, transforming complex business challenges into scalable solutions.
Collaborate with engineering teams, executives, and enterprise clients to define secure, production-ready architectures.
Influence AI adoption strategies, establish best practices, and shape the future of AI-powered solutions.
The company designs and delivers enterprise-grade AI solutions built on advanced cloud technologies. It offers a remote-first culture with a focus on innovation and collaboration among talented engineers.
Design, build, deploy, and maintain production LLM-based solutions and agent workflows.
Lead high-complexity, cross-functional AI initiatives from ambiguous problem definition through production adoption.
Partner with Cybersecurity and Compliance to design PHI-safe data handling patterns and ensure AI solutions align with relevant controls.
Included Health is a healthcare company delivering integrated virtual care and navigation, breaking down barriers to provide high-quality care for every person. They have a remote-first culture and are an Equal Opportunity Employer.