Source Job

India

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
  • Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
  • Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.

Go Python Kubernetes Distributed Systems

20 jobs similar to Senior Software Engineer — Infra Agent Systems

Jobs ranked by similarity.

India

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.

Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.

India

  • Build modular, plug-and-play AI agents that integrate into a broader agentic architecture.
  • Design and implement memory capabilities, including short-term and long-term memory, summarization, and retrieval-backed context.
  • Own quality through testing, evaluations, monitoring, and observability.

They build AI-powered applications to transform proposal workflows in the architecture, engineering, and construction industry. They operate in a fast-moving, remote environment with a focus on experimentation and continuous learning.

$87,480–$110,160/yr
Europe 6w PTO

  • Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
  • Collaborate across teams to improve reliability, deployment complexity, and incident response.
  • Contribute to roadmap planning, technical design, and long-term simplification of stack operations.

Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.

India

  • Build and improve products and infrastructure for AI-powered teams.
  • Contribute across multiple product areas, including CI infrastructure and agentic applications.
  • Take ownership of projects from concept through delivery, with a focus on usability and quality.

Our partner is an early-stage engineering team building infrastructure and tools that help AI-powered teams work more efficiently. It is a high-growth, profitable startup with a small, high-impact team.

US Unlimited PTO

  • Design, architect, and implement AI-powered backend systems that automate and optimize software development workflows.
  • Identify friction points across the engineering lifecycle and transform them into scalable, reliable automation solutions.
  • Collaborate with globally distributed engineering teams to share tooling, best practices, technical roadmaps, and reusable solutions.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

Portugal Poland

  • Architect context pipelines for AI agents, curating what enters the context window so agents produce reliable, mergeable code.
  • Design and maintain agent-ready systems with well-bounded repos, comprehensive test suites, and CI/CD pipelines.
  • Serve as an AI Champion, mentoring engineers on AI-assisted practices and contributing to org-level AI tooling evaluations.

OLX builds marketplace ecosystems enabling millions to buy and sell cars, find homes, land jobs, and trade secondhand goods. Backed by Prosus, the company spans 7 markets with 50+ nationalities and a culture that is ambitious, fast-moving, and built on trust.

$140,000–$225,000/yr
US Canada Unlimited PTO

  • Build and maintain backend services for our LLM gateway, including routing, rate limiting, and observability.
  • Contribute to sandboxing and isolation infrastructure for safe agent-generated code execution.
  • Write high-performance backend code in Go, Rust, or async Python, supporting Kubernetes-based platform services.

AZX accelerates positive impact in critical industries through AI transformation. Founded in 2024 and profitable from the start, we work with category leaders in real estate, energy, logistics, and utilities.

US

  • Design and develop foundational components and frameworks for our Agentic AI platform.
  • Collaborate with cross-functional teams to create platform solutions that empower developers.
  • Provide production support and ensure platform stability, working closely with ML engineers.

Legion builds secure, reliable AI systems that integrate with complex platforms, optimizing workflows and enhancing human capability. They work with partners like Palantir, Nvidia, HPE, and Oracle, and are looking for bold thinkers to shape the future of grounded AI.

India

  • Lead the design and delivery of high-impact prototypes and production-ready solutions using AI and emerging technologies.
  • Mentor engineers and establish strong engineering practices to raise technical standards across the team.
  • Communicate complex technical decisions to executive stakeholders and drive cross-functional collaboration.

This company operates in an emerging technologies and incubation environment focused on building what comes next. They have a globally distributed team and emphasize innovation, autonomy, and mentorship.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

$180,000–$220,000/yr
US Unlimited PTO

  • Design, build, and maintain agent infrastructure and platforms, including the TRACE Graph, embeddings, and semantic search.
  • Collaborate with detection engineers and threat hunters to encode domain expertise into agents.
  • Build automated evaluations and benchmarks for non-deterministic agentic systems.

Nebulock is an agentic threat hunting platform that autonomously surfaces behaviors, not just IOCs, from various data sources. As a startup, we emphasize collaboration, low ego, and a relentless focus on delivering customer value.

$170,000–$190,000/yr
Global Unlimited PTO

  • Design and deploy scalable AI infrastructure and agent systems for enterprise customers.
  • Work on Kubernetes cluster design, multi-agent system architecture, and CI/CD pipelines.
  • Engage directly with customers to assess needs and present technical recommendations.

LangChain builds the foundation for agent engineering, helping developers create production-ready AI agents. With $125M raised at Series B from top venture firms and 100M+ monthly open source downloads, they have a strong engineering culture and meaningful team impact.

$80,000–$120,000/yr
Global Unlimited PTO

  • Design, build, and operate high-load distributed backend services powering the company's ML infrastructure.
  • Take end-to-end ownership of core ML services and data pipelines from design to deployment and continuous improvement.
  • Partner with ML and product teams to understand their needs and turn them into reliable, reusable platform capabilities.

Constructor is an AI-first e-commerce search and discovery platform that helps shoppers find products and enables brands to drive revenue. The company is fully remote, diverse, and values ownership and collaboration.

US UK Ireland Poland Germany Australia

  • Design and implement a scalable observability platform for Whatnot's growing infrastructure.
  • Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
  • Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.

Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.

US

  • Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
  • Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
  • Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.

WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.

Global

  • Design and build the AI execution platform with event-triggered workflows and model-agnostic runtimes.
  • Develop evaluation layers with golden test suites and safety checks to ensure AI reliability.
  • Implement governance mechanisms and build AI agents supporting business workflows.

Jobgether is a platform that connects talent with opportunities using AI-powered matching. The company has a globally distributed team and a remote-first culture.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

India

  • Own work end to end from discovery to production deployment.
  • Design and build data foundations and full-stack applications on top of them.
  • Use AI coding agents as a primary build accelerator to ship production-grade AI systems.

We help companies design, build, and deploy AI agents that automate complex workflows, delivering reliability and efficiency gains. We are a small, fast-moving, ownership-driven team that thrives on building quickly and seeing results in days, not months.