Source Job

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

GPU Computing Kubernetes

20 jobs similar to AI Infrastructure Solutions Architect

Jobs ranked by similarity.

Canada

  • Lead the design and operation of GPU infrastructure for AI workloads.
  • Manage Kubernetes-based environments and optimize for AI training and inference.
  • Define operational standards, implement monitoring, and collaborate with AI engineering teams.

ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.

$180,000–$200,000/yr
Global

  • Lead customers in designing and optimizing GPU-based solutions on Vultr's platform.
  • Collaborate with cross-functional teams to bring AI, ML, and GPU workloads into production.
  • Educate customers on the value of Vultr's cloud infrastructure and expand their possibilities.

Vultr provides high-performance cloud infrastructure solutions globally, making them easy to use, affordable, and locally accessible. It is a privately-held company with over a decade of self-funding, hundreds of thousands of customers across 185 countries, and a culture that emphasizes comprehensive benefits and employee growth.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

$100,000–$150,000/yr
US

  • Design, build, and operate scalable infrastructure platforms for large-scale AI model training and inference.
  • Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
  • Develop software solutions and automation tools using Python and systems programming languages like Go or C++.

Our partner builds and operates foundational technology powering advanced AI training and inference workloads at scale. They offer a collaborative culture focused on innovation, engineering excellence, and continuous learning.

South Korea

  • Lead technical discovery and design scalable AIStor architectures for AI, analytics, and cloud-native applications.
  • Serve as trusted technical advisor to customers, translating business requirements into robust technical designs.
  • Collaborate with cross-functional teams to ensure customer success from evaluation to production deployment.

MinIO is the data and memory foundation for enterprise AI, providing AIStor and MemKV to unify data storage across core, edge, and cloud. They are trusted by 77% of the Fortune 100 and are redefining how AI factories and intelligent applications manage data.

Global Unlimited PTO

  • Partner with go-to-market team to support technical sales and solution design. - Translate customer requirements into workable solutions and build customer confidence. - Create repeatable assets and close the loop with product and engineering.

Massed Compute is building a modern GPU cloud platform for AI and high-performance compute workloads. They are a lean, ambitious team operating in one of the most important technology markets in the world.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

$170,000–$190,000/yr
Global Unlimited PTO

  • Design and deploy scalable AI infrastructure and agent systems for enterprise customers.
  • Work on Kubernetes cluster design, multi-agent system architecture, and CI/CD pipelines.
  • Engage directly with customers to assess needs and present technical recommendations.

LangChain builds the foundation for agent engineering, helping developers create production-ready AI agents. With $125M raised at Series B from top venture firms and 100M+ monthly open source downloads, they have a strong engineering culture and meaningful team impact.

Europe

  • Build and scale massive distributed compute and storage systems for AI training.
  • Architect multi-cluster orchestration and optimize workload placement across global regions.
  • Design future-proof storage formats and implement metadata systems for exabyte-scale growth.

Mistral provides full-stack AI solutions, from frontier models to developer tools. They are a dynamic, collaborative team with a diverse workforce, passionate about innovation and low-ego teamwork.

US 20w maternity 12w paternity

  • Own end-to-end technical execution for strategic customer and partner engagements, including discovery, infrastructure design, implementation, and production deployment.
  • Design and build cloud infrastructure supporting advanced AI workloads, including simulation, training, evaluation, inference, and large-scale batch processing.
  • Improve platform reliability, security, performance, and cost efficiency by debugging issues across application, network, storage, compute, and orchestration layers.

The partner company is building the infrastructure foundation for next-generation AI applications and physical AI workloads. The engineering team is pioneering and values ownership, technical excellence, and solving challenging engineering problems at scale.

India

  • Define and drive innovative technical vision for intelligent networking and agentic AI platforms.
  • Design and build high-performance, production-ready services for real-time data and AI processing.
  • Mentor engineers, lead technical discussions, and uphold engineering excellence across the team.

We provide end-to-end, cloud-driven networking solutions trusted by over 50,000 customers globally. With double-digit growth and a culture of inclusion, we foster an innovative workplace where all employees thrive.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

Global Unlimited PTO

  • Own the technical evaluation end-to-end, from discovery to POC, ensuring evaluations are scoped and tied to customer ROI.
  • Take customers from signature to first successful production training run and serve as the technical owner post-launch.
  • Build the SA function by creating demo environments, benchmarking harnesses, and reference architectures.

Andromeda Cluster gives early-stage startups access to scaled AI infrastructure, partnering with leading AI labs and cloud providers. It is a high-growth company building an inclusive environment for all employees.