Source Job

$220,000–$300,000/yr

  • Build production runtime infrastructure in Rust and Go for Modal's container platform.
  • Optimize container startup, checkpoint/restore, and snapshot pipelines for multi-GPU inference.
  • Diagnose complex Linux, kernel, and performance issues across the stack.

Rust Go Linux GPU

16 jobs similar to Member of Technical Staff - Inference Runtime

Jobs ranked by similarity.

China

  • Operate and expand Telnyx's own B300 GPU fleet to maximize inference throughput per GPU-dollar.
  • Design and implement serverless inference for open-weight models and dedicated enterprise deployments.
  • Work upstream in open-source technologies like vLLM, SGLang, and Kubernetes.

Telnyx is an industry leader building the future of global connectivity through a private, multi-cloud IP network and edge APIs. The company is financially stable and profitable, with a global team and a focus on innovation and continuous learning.

Canada

  • Lead the architecture and delivery of a large-scale GPU infrastructure platform, evolving from managed Kubernetes to bare-metal with Slurm and inference support.
  • Manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting technical standards and overseeing implementation.
  • Own GPU infrastructure operations, including Slurm, Kubernetes, NVIDIA hardware, observability, and incident response, while acting as the primary technical interface with partners.

Jobgether is an AI-powered job matching platform that connects candidates with relevant roles, ensuring a fair and objective review process. It operates with a distributed team and partners with companies globally, focusing on efficient and transparent recruitment.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

$150,000–$220,000/yr
Global Unlimited PTO

  • Lead the effort to make Runpod the fastest and most cost-efficient place for LLM inference, owning performance end to end.
  • Profile and diagnose performance bottlenecks across the serving stack, from scheduling to kernels, and implement fixes.
  • Work closely with product and infrastructure teams to shape how inference is offered, turning improvements into production-ready defaults.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We're a small, remote-first team that takes ownership seriously, moves fast, and has processed more than 20 billion inference requests.

Global

  • Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
  • Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
  • Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.

Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.

$205,000–$231,000/yr
US Canada Unlimited PTO 18w maternity 12w paternity

  • Drive technical vision and set the roadmap for the squad, partnering with engineering leadership.
  • Architect the end-to-end automation platform for package creation, test generation, and image building.
  • Build AI-powered tooling with LLM integration for manifest generation and quality gates.

Chainguard delivers hardened, secure builds of open source software to secure the software supply chain. They serve Fortune 500 enterprises and are backed by top investors, with a remote-first culture focused on trust, intentional action, and customer obsession.

India

  • Build and maintain fleet-wide OTA update capabilities for thousands of on-premise appliances.
  • Develop and optimize control-plane connectivity between deployed appliances and AWS infrastructure.
  • Tune network stack performance and conduct profiling to improve reliability across a distributed fleet.

They build and operate core edge infrastructure deployed across thousands of on-premise appliances in school networks. The company offers a remote senior systems engineering role, with a distributed team collaborating across India, the US, and the UK.

$184,100–$263,000/yr
North America

  • Own Docker Sandboxes, including its long-term vision, roadmap, revenue, adoption, and results.
  • Engage deeply with developers building AI agents, coding environments, CI and testing systems, and workloads requiring secure, isolated execution.
  • Turn complex tradeoffs across isolation, startup time, performance, density, networking, storage, and cost into simple product experiences.

Docker builds developer tools for building, sharing, and running applications, including Docker Desktop, Docker Hub, and Docker Scout. It is a globally distributed, remote-first team trusted by over 20 million monthly users and 20 billion container image pulls.

  • Build Go services on the event bus, AI/ML pipelines, or connector frameworks for a high-throughput, low-latency platform.
  • Own your latency budget and committed timelines in a fast startup environment with 1-week sprints and Friday demos.
  • Demonstrate ideas easily with working prototypes, adapt quickly to new technology, and ship clean code fast.

TTEC Digital is an innovation group inside TTEC, building next-generation AI CX tools for the world's biggest brands. As a public company since 1982, they operate like an early-stage startup with enterprise scale and a culture that values employee growth and innovation.

$100,000–$300,000/yr
US Unlimited PTO

  • Lead the architecture and development of the Hercules Database, backend services, and SDKs.
  • Engineer for low latency, high throughput, and efficient resource use across data-intensive, multi-tenant systems.
  • Build and operate reliable hosting infrastructure with strong tenant isolation, observability, and failure resilience.

Hercules builds the future of AI software development with a platform-as-a-service including a database, backend services, and SDKs. It is an early-stage company with a founding team based in San Francisco and a culture that values shipping fast, excellence, and hunger.

Kenya

  • Design and build production-grade systems end-to-end, from problem definition through deployment and operations.
  • Work across application services, distributed systems, infrastructure, data pipelines, and ML systems, debugging complex issues across multiple layers.
  • Frame problems correctly, applying ML when needed, and ensure reliability, performance, and cost efficiency.

Moniepoint Inc. is Africa's all-in-one financial platform, helping 20 million businesses and individuals access payments, banking, credit, cross-border, and business management tools. As Nigeria's largest merchant acquirer processing over $250 billion annually, we prioritize our people's well-being and foster a culture of innovation and teamwork.

$157,000–$184,000/yr
US Canada Unlimited PTO 18w maternity 12w paternity

  • Design and build automation systems for package creation, test generation, and image building.
  • Develop AI-powered tooling using LLMs for manifest generation and validation.
  • Write production Go code and create quality tools to improve customer reliability.

Chainguard delivers hardened, secure, and production-ready builds of open source software. Backed by leading investors, they serve Fortune 500 enterprises and global industry leaders.

$192,000–$192,000/yr
Global

  • Lead large-scale Sensor Platform initiatives in collaboration with product, infrastructure, and engineering peers.
  • Own root cause analysis and postmortems for complex production issues, ensuring durable fixes.
  • Establish standards for documentation, test harnesses, and observability across the codebase.

Dragos is the global leader in OT cybersecurity, combining technology, threat intelligence, and expert services to protect critical infrastructure. The team is remote-first, mission-driven, and built on authenticity, transparency, and trust.

Europe

  • Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
  • Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
  • Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.

Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.

$90,000–$110,000/yr
US EMEA

  • Act as senior technical resource and final escalation point for strategic and VIP customers, owning complex issues across Kubernetes, GPU, and enterprise stack.
  • Train and mentor Technical Support Engineers in advanced Linux troubleshooting and customer architectures.
  • Author advanced troubleshooting documentation and drive incident resolution through root cause analysis.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. It is the world's largest privately-held cloud infrastructure company with 33 data centers and hundreds of thousands of active customers, committed to growth and employee investment.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.