Source Job

Global

  • Design and build a managed Slurm service on Kubernetes
  • Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
  • Build observability and automated remediation for GPU, node, network, and control-plane failures

Go Kubernetes GPU Distributed Systems

20 jobs similar to Software Golang Engineer (Slurm)

Jobs ranked by similarity.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.

US Europe

  • Design and build control-plane services, drivers, and tooling for high-performance storage integration with Kubernetes.
  • Write production-quality Go software with strong testing and operational rigor for automation.
  • Deliver storage integration for Kubernetes via Cluster API and K0rdent in hybrid and air-gapped environments.

Mirantis is a Kubernetes-native AI infrastructure company enabling organizations to build scalable infrastructure for AI and data-intensive applications. A Silicon Valley leader with a young, collaborative culture.

US

  • Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
  • Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
  • Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.

Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.

APAC

  • Design, build, and operate components of the Kubernetes platform and its core subsystems end to end.
  • Write and review Go code for controllers, operators, platform services, and automation.
  • Help operate the platform, including on-call, incident investigation, and follow-up work to prevent recurrence.

ServiceNow provides an AI platform for business reinvention, helping 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent collaborate.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.

US

  • Design and build control-plane services and drivers for storage integration with Kubernetes-based AI workloads.
  • Write production-quality Go code with strong testing and operational rigor.
  • Deliver storage integration for k0s-based Kubernetes via Cluster API and K0rdent topologies.

Mirantis is a Kubernetes-native AI infrastructure company that builds scalable, secure infrastructure for AI and data-intensive applications. It is committed to open standards and freedom from lock-in, empowering platform engineering teams.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

US

  • Dive deep into Kubernetes and related technologies like Istio, Helm, and Prometheus.
  • Tackle daily engineering challenges and drive innovation in Kubernetes optimization.
  • Work with cutting-edge tools and frameworks such as container orchestration, microservices architecture, and cloud-native applications.

We are redefining autonomous cloud and AI infrastructure, freeing DevOps teams from manual resource management to maximize performance and reduce cloud costs by up to 80%. Backed by over $210M from leading VCs and trusted by Fortune 100 companies, we are building a world-class platform with a culture of innovation.

$189,000–$231,000/yr
North America Unlimited PTO

  • Own the design and implementation of platform subsystems end to end.
  • Anchor meaningful projects, balancing reliability, scalability, and time to market.
  • Mentor SWE and SWE II engineers and contribute to a strong code review culture.

Stacklok is building the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on their own infrastructure. The company is led by Kubernetes co-creators and is a startup with a collaborative, AI-maximalist culture.

Germany Unlimited PTO

  • Develop and maintain scalable Go code for AI/ML components within a cloud platform.
  • Design and improve AI/ML services ensuring reliability, performance, and maintainability.
  • Collaborate with cross-functional teams to translate requirements into technical solutions.

Jobgether uses AI-powered matching to connect candidates with roles. They are a remote-first company with a focus on efficient and fair hiring processes.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

US

  • Build and maintain SRE microservices under the guidance of senior engineers.
  • Deliver infrastructure and application updates using GitOps and automated CI/CD pipelines.
  • Participate in a mentored on-call rotation and write clear runbooks and incident post-mortems.

Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure solutions. Headquartered in Singapore, the company operates data centers across multiple countries and focuses on building computational infrastructure.

US Unlimited PTO 18w maternity 12w paternity

  • Collaborate on the Developer Platform to improve CI/CD speed, reliability, and developer experience.
  • Help productionize an Agentic AI platform with focus on observability and context-engineering.
  • Partner with engineering teams to consolidate build systems and automate repetitive manual work.

Chainguard delivers hardened, secure, and production-ready builds of open source software, helping organizations build faster and eliminate risk. The company is venture-backed by leading investors including Sequoia Capital, and serves Fortune 500 enterprises.

US Unlimited PTO 18w maternity 12w paternity

  • Build, operate, and improve Elastic Build, a Kubernetes-based pipeline for producing production-ready artifacts.
  • Design and implement automation for package rebuild and review workflows while maintaining human oversight.
  • Develop monitoring, dashboards, alerting, and automated remediation to reduce operational toil and detect issues early.

Chainguard provides hardened, secure, and production-ready builds of open source software. The company is venture-backed by leading investors and serves Fortune 500 enterprises and global industry leaders, fostering a remote-first culture with values of customer obsession, intentional action, and trust.

North America Unlimited PTO

  • Deploy and scale MCP-based AI agents on Kubernetes for enterprise customers across the US-East and EMEA regions.
  • Lead complex technical engagements, build reusable deployment patterns, and mentor engineers on the team.
  • Shape product roadmap by feeding back field insights from regulated industries and defining regional engagement standards.

Stacklok builds the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on Kubernetes and private cloud. Founded by two Kubernetes creators, the company is already adopted by leading tech and regulated industries, fostering a collaborative, AI-maximalist culture with deep open-source roots.

$182,000–$238,000/yr
US

  • Design, code, and debug scalable software applications and APIs in Golang with attention to performance and security.
  • Collaborate with product managers and cross-functional teams to translate business requirements into technical specifications.
  • Provide technical leadership, drive code reviews, and guide the team in AI tooling best practices.

EasyPost is a YC unicorn that makes shipping simple for businesses with the first developer-friendly REST API for shipping. Our team is rapidly growing, and we're builders and problem-solvers who move fast and innovate in an industry that needs it.

$130,000–$145,000/yr
US

  • Build and maintain high-performance backend services and cloud-native applications using Go.
  • Design and optimize REST and gRPC APIs for scalability, performance, and reliability.
  • Collaborate with cross-functional teams to translate requirements into scalable software solutions.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses technology to review applications and share top-fitting candidates with employers, focusing on efficient and objective recruitment.

United States

  • Design and develop high-performance backend services using Go.
  • Architect scalable distributed systems and microservices with a focus on reliability and observability.
  • Mentor engineers, guide architectural decisions, and resolve complex production challenges.