Accountabilities:
- Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments.
- Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents.
- Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support.
Requirements:
- Bachelor’s degree or equivalent experience in Computer Science or related field, with 5+ years building production backend or distributed systems.
- Deep expertise in AI agent systems, orchestration, knowledge graphs, or retrieval-augmented generation.
- Proficiency in Go, TypeScript, Python, or Rust, and experience with Kubernetes, GitOps, and infrastructure-as-code.
Benefits:
- Fully remote position based in India with opportunity to work on AI agents at significant infrastructure scale.
- End-to-end ownership of production software from design through deployment and operations.
- Collaborative environment with technically ambitious engineers and researchers tackling challenging AI infrastructure problems.
Partner Company
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.