Similar Jobs

See all

Role Overview:

  • Own and shape critical ML infrastructure for a rapidly scaling AI platform.
  • Design reliable, production-grade machine learning platforms supporting large-scale AI workloads.
  • Build foundational systems from the ground up in a remote-first environment.

Key Responsibilities:

  • Optimize GPU utilization, memory efficiency, and network performance.
  • Develop scalable deployment pipelines using blue/green and canary rollouts.
  • Implement observability solutions for inference latency, throughput, and cost.

Technical Requirements:

  • 4+ years in MLOps, Platform Engineering, or SRE roles.
  • Strong experience with vLLM, TGI, or Triton in production.
  • Proficiency in Python, Terraform, and containerized environments.

Jobgether

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

Apply for This Position