Source Job

Latin America

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies, and safe rollback.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, guardrails, and graceful degradation under load.

Terraform Kubernetes GCP CI/CD Machine Learning

20 jobs similar to Platform Architect

Jobs ranked by similarity.

UK Germany Netherlands Ireland Spain Poland Bulgaria Lithuania Unlimited PTO

  • Build and own the model serving infrastructure, real-time inference, feature retrieval, and the latency budget that governs both.
  • Build the deployment path for data scientists to ship models, including bring-your-own-model support.
  • Own models in production: monitoring, drift detection, retraining, incident response, and the on-call rotation.

Sardine is the leading agentic risk platform for fighting financial crime. We are a remote-first company with hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo, hiring talented individuals with extreme ownership and high growth orientation.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

US Unlimited PTO

  • Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
  • Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
  • Oversee the sandbox runtime, control-plane services, and observability tooling.

AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

$100,000–$145,000/yr
US

  • Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
  • Lead infrastructure automation and release management, driving best practices in security and compliance.
  • Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.

Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.

Latin America

  • Build AI-powered tools and copilots across the SDLC to reduce cognitive load and eliminate manual steps.
  • Research and deploy GenAI solutions to improve delivery pipelines and system reliability.
  • Collaborate with Platform, SRE, and DevOps teams to integrate intelligent automation into the core engineering platform.

Coderio designs and delivers scalable digital solutions for global companies. They combine strong technical expertise with a product mindset and value autonomy and clear communication.

Global 16w maternity 16w paternity

  • Develop tools and automate manual processes to improve operational efficiency and accelerate experimentation.
  • Build, integrate, and monitor end-to-end lifecycles of large-scale distributed machine learning systems.
  • Enhance ML pipelines for forecasting platforms with automated retraining and deployment.

Apella applies computer vision and machine learning to improve the standard of care in surgery. They are a remote-first startup focused on building reliable ML solutions for healthcare.

$149,200–$214,500/yr
US

  • Architect, design, build, deploy, and maintain Model Serving infrastructure using industry-standard AI tools.
  • Own projects that scale model serving and data processing services to handle 10x traffic.
  • Collaborate closely with MLE and Data Science teams to distill feedback and execute on strategy.

Abnormal AI protects the humans behind the world's most critical organizations from AI-powered cybercrime. Over 4,500 enterprises trust their behavioral AI platform, fostering a culture of innovation and security.

$161,000–$221,500/yr
United States

  • Drive AI platform architecture and execute the long-term roadmap for machine learning and generative AI.
  • Lead AI infrastructure vision including end-to-end training, fine-tuning, and low-latency inference platforms.
  • Set engineering excellence standards for the full AI/ML SDLC and mentor senior engineers.

Lyra Health is a leading provider of evidence-based mental health care, serving over 20 million people globally. The company operates with a focus on transformative care and has delivered over 15 million sessions, employing a large team that includes engineers, data scientists, and clinical professionals.

Romania

  • Architect, deploy, and manage highly available, fault-tolerant cloud infrastructure across Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
  • Maintain and scale declarative infrastructure using Terraform across a multi-hundred-file estate, enforcing GitOps workflows with Atlantis.
  • Build, maintain, and optimize robust automated pipelines for continuous integration and delivery using GitHub Actions, Jenkins, and ArgoCD.

Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is a scrappy, nimble organization dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.

$250,000–$285,000/yr
US Unlimited PTO

  • Define architecture and best practices for the platform and infrastructure layer the product is built on.
  • Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
  • Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.

We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.

US Unlimited PTO

  • Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
  • Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
  • Own platform reliability, observability, capacity planning, and operational readiness.

OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

Romania

  • Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
  • Automate operations and engineering tasks to improve productivity and reliability.
  • Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.

Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.

Poland Unlimited PTO

  • Own the ML/AI platform including training infrastructure, model serving, inference pipelines, and production integration.
  • Drive the GenAI/LLM strategy including retrieval architectures, evaluation harnesses, and agentic workflows.
  • Partner with data scientists, DevOps, and backend engineers to productionize models and define API contracts.

SavvyMoney is a US-based financial technology company providing integrated credit score and personal finance solutions to over 1,600 banks and credit unions throughout the United States. The company was recognized as a 'Top 25 Places to Work' in the San Francisco Bay Area and is an Inc. 5000 Fastest Growing Company.

US

  • Design and maintain CI/CD and MLOps pipelines for AI and software applications, ensuring seamless deployment and automation.
  • Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters to support high-performance AI workloads.
  • Champion Infrastructure as Code and observability practices to ensure high availability, security, and compliance across multi-cloud environments.

Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company has a global presence with data centers in multiple countries and a culture focused on innovation and reliability.

Global

  • Own the technical path from customer interest to working deployment, integrating the platform into production AI environments.
  • Build and operate AI/MLOps pipelines, debug complex environments, and create prototypes and demos.
  • Translate customer needs into product improvements, partnering with Sales, Product, and Engineering.

Neuromorphic Labs is a Seed-stage AI startup building a trust layer for production AI, making security, governance, and control intrinsic to every model and deployment. Backed by top-tier VCs, the team is small and fast-paced, emphasizing ownership, high standards, and collaboration with exceptional builders.

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

$140,000–$175,000/yr
US

  • Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
  • Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
  • Partner cross-functionally to align platform priorities with business goals and ensure system reliability.

PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.