Source Job

$220,000–$320,000/yr
US

  • Own and improve the platform that processes customer jobs, focusing on speed, reliability, and resource efficiency.
  • Partner with machine learning teams to serve models effectively and integrate them into production workflows.
  • Build and maintain infrastructure for databases, task orchestration, workload scheduling, and cloud deployments.

Python Go Rust Distributed Systems Cloud Infrastructure

20 jobs similar to Platform Engineer

Jobs ranked by similarity.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

Spain

  • Design and develop large-scale platforms for LLM training and AI workloads.
  • Tackle distributed-systems challenges including intelligent job scheduling and resource optimization.
  • Collaborate with international teams to build production-ready AI infrastructure.

This role is with a partner company, an AI-focused R&D team building infrastructure for large language models. They are a fast-moving, highly technical team with a collaborative and innovative culture.

$190,000–$230,000/yr
US

  • Design and implement scalable cloud infrastructure using Kubernetes, Pub/Sub, and distributed systems technologies.
  • Collaborate with our AI team to optimize data pipelines and integrate AI to remove performance bottlenecks.
  • Drive platform reliability initiatives including alerting, health checking, and incident management.

Syllo is building a unified litigation platform that helps lawyers and paralegals use AI throughout the litigation life cycle. We are a quickly expanding company with enterprise customers including major law firms and corporations.

India

  • Own and evolve distributed systems architecture for platform services, ensuring reliability and performance at enterprise scale.
  • Drive cost optimization, observability, and incident management practices across cloud infrastructure.
  • Hire, mentor, and grow engineers while partnering cross-functionally to align platform capabilities with organizational needs.

HighLevel is an AI-powered business operating system that helps agencies, entrepreneurs and SMBs build, automate and scale their operations. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

$160,000–$250,000/yr
US

  • Own critical infrastructure components and make architectural decisions with significant autonomy.
  • Build and improve reliable, secure, scalable systems for enterprise customers.
  • Develop tools, systems, and engineering practices that help product teams ship quickly.

US Unlimited PTO

  • Lead the architecture and development of the Hercules Database, backend services, and SDKs.
  • Engineer for low latency, high throughput, and efficient resource use across data-intensive, multi-tenant systems.
  • Build and operate reliable hosting infrastructure with strong tenant isolation, observability, and failure resilience.

Hercules builds the future of AI software development with a platform-as-a-service including a database, backend services, and SDKs. It is an early-stage company with a founding team based in San Francisco and a culture that values shipping fast, excellence, and hunger.

US

  • Build core data platform capabilities, developing services and APIs for data access, query execution, ingestion, and synchronization.
  • Make integrations reusable by building connector frameworks that support different systems without bespoke implementations.
  • Own execution reliability by building orchestration, retries, checkpointing, and observability for predictable workload recovery.

Tessera Labs is redefining how enterprises adopt and operationalize AI by building multi-agent systems that automate complex business workflows across platforms like SAP and Salesforce. Backed by top venture capital firms and built by leaders from Meta AI and Google Research, they move fast and operate with extreme ownership.

$145,500–$235,400/yr
US

  • Design, implement, test, and operate production services and APIs.
  • Lead projects or meaningful components of projects from problem definition through deployment and iteration.
  • Improve the observability and operability of the systems you own, including metrics, logs, traces, alerting, and incident learnings.

LaunchDarkly provides a platform that helps engineering teams release software and AI with speed, safety, and control using feature flags and observability. The company is growing and emphasizes teamwork, humility, openness, curiosity, and inclusive collaboration.

Canada

  • Architect, design, build, deploy, and maintain Model Serving infrastructure for a world-class Detection Engine.
  • Own projects that scale model serving and data processing to handle 10x traffic, including real-time streaming pipelines and online feature serving.
  • Collaborate with MLE and Data Science teams to build the ML Training platform, improving MLE velocity and model precision and recall.

Abnormal protects the humans behind the world's most critical organizations from AI-powered cybercrime. 4,500+ enterprises trust our behavioral AI platform, and we foster a culture of innovation and impact.

  • Architect AI infrastructure including backend services, data pipelines, and orchestration workflows for LLMs.
  • Write production-grade, high-performance code for high-throughput AI workflows and data-driven systems.
  • Integrate foundational models, vector data stores, and cloud microservices into scalable platform components.

Korza builds AI infrastructure and products, specializing in LLM and agent-based systems. The company emphasizes cross-functional collaboration and rapid shipping of cutting-edge AI capabilities.

North America Unlimited PTO

  • Take deep ownership of systems and product areas, becoming the go-to technical expert for what you own.
  • Architect, build, and maintain highly available, scalable REST APIs and backend services across distributed systems.
  • Build agentically with AI coding tools, staying at the forefront of emerging AI tooling and practices.

Postscript is an AI messaging platform for Shopify brands, helping e-commerce marketers drive revenue through SMS marketing and AI features. Backed by Greylock and Y Combinator, the company has been fully remote since 2018 and serves 20,000+ Shopify brands with a culture of ownership and AI-first development.

US

  • Design and build full-stack capabilities supporting AI initiatives across Keeper's product portfolio.
  • Develop backend services, APIs, and frontend experiences that integrate AI-driven functionality into secure production environments.
  • Collaborate with AI and product engineering teams to remove technical dependencies and deliver production-ready AI features.

Keeper Security is a cybersecurity software company providing zero-knowledge and zero-trust security solutions that protect passwords, secrets, and infrastructure. It is a fast-growing global organization trusted by millions of users and thousands of enterprises, with a culture focused on innovation and reliability.

$97,600–$139,000/yr
United States Canada

  • Build and maintain core infrastructure for Quora's ML platform, ensuring high availability, scalability, and performance.
  • Build and improve distributed systems serving ML models in production, from Large Recommendation Models to Large Language Models.
  • Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.

Quora's mission is to grow the world's collective intelligence through two platforms: Quora for global knowledge sharing and Poe for AI agent collaboration. We are a remote-first company with passionate, collaborative, and high-performing global teams, rooted in transparency and experimentation.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

UK

  • Design and ship backend features that shape a global platform, making sound architectural trade-offs.
  • Mentor engineers, define standards, and drive quality, security, and operational excellence.
  • Integrate agentic AI into production workflows and collaborate with PMs and designers to deliver user value.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity and accelerate digital transformation. Trusted by more than 50 million registered users and over half of the Fortune 100, GitLab fosters a remote-first, AI-driven culture where every voice is valued.

Global

  • Build LLM-based agents on the platform's scaffolding, integrating tool calls, internal APIs, and guardrails.
  • Take agents to production on AWS with containers, CI/CD, secrets, permissions, and security controls.
  • Define and run evals, monitor with Langfuse, and document runbooks for independent operation.

Muttdata builds innovative Data Products and Machine Learning solutions to help companies solve complex business challenges. It is a fast-growing, remote-first startup that values collaboration, continuous learning, and a positive, ownership-driven culture.

North America

  • Take deep ownership of key systems and product areas, becoming the technical expert and primary point of reference.
  • Architect, build, and maintain highly available and scalable REST APIs and backend services.
  • Use agentic development workflows and AI coding tools as standard engineering practice.

This partner company builds an AI-driven e-commerce messaging platform powered by customer context and data. It is a high-growth, remote-first environment with a collaborative and supportive culture.

Global

  • Own features end-to-end: prototype, architect, build, test, deploy, and support production services.
  • Build and scale backend services and APIs in Python powering Monte Carlo's agent trust platform.
  • Design distributed systems and data pipelines while shipping React frontends where needed.

Monte Carlo is the agent trust platform that unifies data and agent observability to monitor, troubleshoot, and improve production AI systems. Founded in 2019 and backed by leading investors, the company is remote-first and recognized as a best workplace.