Source Job

China

  • Operate and expand Telnyx's own B300 GPU fleet to maximize inference throughput per GPU-dollar.
  • Design and implement serverless inference for open-weight models and dedicated enterprise deployments.
  • Work upstream in open-source technologies like vLLM, SGLang, and Kubernetes.

Kubernetes GPU

20 jobs similar to Inference Infrastructure Architect

Jobs ranked by similarity.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.

US Unlimited PTO

  • Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
  • Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
  • Oversee the sandbox runtime, control-plane services, and observability tooling.

AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.

Europe

  • Lead in-depth technical discovery with engineering teams and customer stakeholders to understand AI inference requirements.
  • Translate customer objectives into production-ready architectures and define PoC success criteria.
  • Identify recurring workload patterns and communicate insights to Product and Engineering for platform evolution.

They are a partner company focused on AI infrastructure and performance-sensitive AI inference workloads. They have an international, engineering-led team solving complex challenges at the forefront of AI.

Global

  • Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
  • Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
  • Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.

Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.

Canada

  • Lead the architecture and delivery of a large-scale GPU infrastructure platform, evolving from managed Kubernetes to bare-metal with Slurm and inference support.
  • Manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting technical standards and overseeing implementation.
  • Own GPU infrastructure operations, including Slurm, Kubernetes, NVIDIA hardware, observability, and incident response, while acting as the primary technical interface with partners.

Jobgether is an AI-powered job matching platform that connects candidates with relevant roles, ensuring a fair and objective review process. It operates with a distributed team and partners with companies globally, focusing on efficient and transparent recruitment.

India

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.

Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.

US

  • Architect and own the technical roadmap for a secure local LLM platform deployed in company-controlled infrastructure.
  • Build modular inference layers with stable APIs, model routing, and production-grade serving optimizations.
  • Design and operate retrieval-augmented generation pipelines with permission-aware access and systematic evaluation.

Parallel Wireless is a U.S.-based pioneer in Open RAN innovation, transforming how mobile networks are built and powered. The company is a leader in software-centric, hardware-agnostic network solutions with a focus on reducing complexity and total cost of ownership.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.

Europe

  • Deploy and commission GPU cloud infrastructure across regional and core datacenters, covering hardware, networking, storage, and platform software.
  • Act as a hands-on technical escalation point, troubleshooting complex issues across physical and software layers and driving incidents to resolution.
  • Establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.

The company builds and operates cutting-edge GPU cloud and AI infrastructure at significant scale. It is a fast-growing international scale-up with a diverse and flexible working environment.

Latin America

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies, and safe rollback.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, guardrails, and graceful degradation under load.

ReadyOn is an AI-native Labor Operating System that redefines how enterprises manage frontline labor by matching workers to shifts in real time. Headquartered in San Francisco with over 100 employees, it grew revenue 8x year over year in 2025.

Global

  • Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
  • Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
  • Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.

The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.

$200,000–$230,000/yr
Global

  • Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
  • Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
  • Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.

Europe

  • Own the technical relationship and be the trusted advisor for CTOs and platform teams.
  • Architect real solutions that translate customer constraints into deployable AI infrastructure.
  • Lead proofs of concept under real conditions to demonstrate operational fit.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable and secure infrastructure for modern AI workloads. The company has an installed base of 1,500 enterprise customers and values open source innovation, collaboration, and continuous growth.

North America Unlimited PTO

  • Build reference architectures, benchmarks, and documentation that engineers trust.
  • Own the developer community, answering hard questions and setting the tone.
  • Write production-grade code for integrations and tooling that lower the barrier for new users.

Andromeda Cluster gives early-stage startups access to scaled AI infrastructure once reserved for hyperscalers. We are a unicorn with a small senior team, building the liquidity layer for global AI compute.

US

  • Define the strategy, roadmap, and feature priorities for k0rdent AI Kubernetes services, empowering Neocloud operators to launch managed Kubernetes on their own GPU infrastructure.
  • Translate requirements from NeoClouds, GPU clouds, telcos, and enterprise platform teams into clear product direction and partner with engineering to ship secure, scalable cluster lifecycle capabilities.
  • Manage the Kubernetes backlog, define positioning and competitive differentiation, and create field-facing assets to support strategic accounts.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They have a world-class, distributed team committed to openness and technical excellence.

$140,000–$225,000/yr
US Canada Unlimited PTO

  • Build and maintain backend services for our LLM gateway, including routing, rate limiting, and observability.
  • Contribute to sandboxing and isolation infrastructure for safe agent-generated code execution.
  • Write high-performance backend code in Go, Rust, or async Python, supporting Kubernetes-based platform services.

AZX accelerates positive impact in critical industries through AI transformation. Founded in 2024 and profitable from the start, we work with category leaders in real estate, energy, logistics, and utilities.

US

  • Deploy and integrate high-performance NFS-based storage into Kubernetes clusters via CSI for GPU-accelerated workloads.
  • Automate storage provisioning and monitoring using infrastructure-as-code tools like Terraform and GitOps pipelines.
  • Tune Linux and network settings to optimize throughput and latency for demanding AI and machine learning applications.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable, secure infrastructure for AI and data-intensive workloads. With deep expertise in open source and Kubernetes orchestration, they enable platform engineering teams across on-premises, cloud, edge, and sovereign environments.

India

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
  • Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
  • Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.

The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.

Asia

  • Build and improve core components of the Agent Runtime, including intent routing, RAG, and multi-step agent workflows.
  • Design and optimize Tool Server / Tool Calling capabilities for agent execution and integration.
  • Develop evaluation datasets and harnesses to measure agent quality and reliability.

Binance is a leading global blockchain ecosystem behind the world's largest cryptocurrency exchange by trading volume and registered users. With over 320 million users in 100+ countries, Binance offers a dynamic, flat-structured environment focused on innovation and growth.