Source Job

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Python Kubernetes Terraform AWS PyTorch

20 jobs similar to Senior AI Platform Engineer (Enterprise Systems)

Jobs ranked by similarity.

Latin America

  • Build AI-powered tools and copilots across the SDLC to reduce cognitive load and eliminate manual steps.
  • Research and deploy GenAI solutions to improve delivery pipelines and system reliability.
  • Collaborate with Platform, SRE, and DevOps teams to integrate intelligent automation into the core engineering platform.

Coderio designs and delivers scalable digital solutions for global companies. They combine strong technical expertise with a product mindset and value autonomy and clear communication.

Philippines

  • Participate in a structured training track to develop expertise in AI benchmarking, profiling, and performance tuning.
  • Build and scale benchmarking infrastructure for evaluating AI systems in enterprise settings.
  • Design agent evaluation pipelines that measure reasoning, accuracy, alignment, and user outcomes.

DevRev is building Computer, an AI teammate that unifies data, tools, and workflows into a single AI-ready platform, giving employees real-time insights and proactive suggestions. Backed by Khosla Ventures and Mayfield with over $150M raised, the company is trusted by global companies across industries and fosters a culture of innovation and collaboration.

US

  • Design, build, and improve production agentic AI systems used to solve complex real-world problems.
  • Develop agent architectures, model integrations, and evaluation frameworks for reliable, production-grade AI.
  • Build scalable APIs, services, and infrastructure supporting agent execution and AI-powered product experiences.

Air is the leader in Enterprise Readiness, providing an AI-native platform to align development, production, delivery, and sustainment for government agencies and industrial suppliers. The company is a startup with a focus on mission-driven work and innovation.

US Unlimited PTO

  • Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
  • Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
  • Oversee the sandbox runtime, control-plane services, and observability tooling.

AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.

India

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.

Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.

$170,000–$190,000/yr
US Unlimited PTO

  • Design and own LLM-powered agents end-to-end on AWS GenAI platform with LangChain/LangGraph and Bedrock.
  • Build and operate MCP servers that expose company data and services as structured tools for AI models.
  • Define testing, reliability, and operational practices for major agent services and mentor across the team.

Mitratech builds world-class products that simplify operations in Legal, Risk, Compliance, and HR functions. They serve 20,000 client companies, including 30% of the Fortune 500, and have a diverse, inclusive culture that supports individual excellence.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

Colombia 1w PTO

  • Operate as a technical leader shaping architecture and quality, mentoring developers, and translating technical decisions into business value.
  • Design, build, and operate agentic AI systems including LLM services, retrieval pipelines, and multi-agent orchestration with end-to-end accountability.
  • Lead AI-augmented development practices, evaluating emerging technologies and establishing team-level guardrails for safe and effective AI tool use.

Caseware is a Canadian fintech company that has led the global audit and accounting software industry for over 30 years, serving over 500,000 users across 130 countries. With more than 36,000 professionals listing Caseware as a skill, they foster a culture of independence, innovation, trust, and accountability in a growing global SaaS environment.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

US

  • Drive backend development for AI workflows using Python and FastAPI as part of a collaborative team.
  • Productionize LLM integrations with systems for Bedrock usage, quotas, retries, and cost controls.
  • Build for scale by optimizing async job orchestration, performance, and data-layer for petabyte-scale enterprise data.

Smarsh empowers organizations to manage risk and uncover intelligence in digital communications. With over 6500 clients in regulated industries and consistent recognition from Gartner and Forrester, Smarsh has been listed on the Inc. 5000 as one of America's fastest-growing companies since 2008.

US Unlimited PTO

  • Architect and deploy autonomous AI agents and multi-agent workflows for privacy-first systems.
  • Build scalable backend services using FastAPI and orchestrate agentic workflows with LangGraph in AWS or Azure.
  • Develop rigorous evaluation pipelines for accuracy, citation adherence, latency, and reliability.

Osano is a leading data privacy platform that helps organizations comply with global privacy regulations like GDPR and CCPA. Backed by top-tier investors and recognized as a Great Place to Work for four years running, the company has a 97% employee satisfaction rate and a mission-driven, fast-growing culture.

Latin America

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies, and safe rollback.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, guardrails, and graceful degradation under load.

ReadyOn is an AI-native Labor Operating System that redefines how enterprises manage frontline labor by matching workers to shifts in real time. Headquartered in San Francisco with over 100 employees, it grew revenue 8x year over year in 2025.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.

US

  • Design and build sandboxed evaluation environments for AI models to safely execute code and interact with tools.
  • Build backend services and infrastructure supporting large-scale AI and agentic evaluations.
  • Develop agent scaffolding, evaluation harnesses, and systems for provisioning isolated environments using Docker, Kubernetes, and cloud infrastructure.

10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. The company focuses on adversarial red teaming and model evaluations to help teams deploy AI systems safely.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

$200,000–$350,000/yr
US

  • Develop and deploy machine learning and AI systems.
  • Work with LLMs, generative AI, and modern ML frameworks.
  • Optimize model performance, latency, and cost.

A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.

India

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
  • Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
  • Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.

The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

$170,170–$286,000/yr
North America

  • Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
  • Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
  • Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.

Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.