Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Design, build, and operate infrastructure for real-time systems handling millions of concurrent connections and billions of monthly API requests.
Drive Kubernetes end to end: cluster architecture, workload design, and migration of existing services from AWS to GCP.
Own cloud cost and efficiency optimization, measuring impact against real spend and utilization data.
Stream powers real-time chat, video, activity feeds, and AI moderation for billions of end-users across thousands of apps. We are a Series B company with around 145 employees from over 35 countries, offering a fast-paced startup culture with real ownership.
Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
Lead incident response and implement observability to improve availability and performance.
Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.
Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.
Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.
Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Own best practices for managing production infrastructure, including provisioning, scaling, configuration, capacity planning, and monitoring.
Build and maintain Kubernetes infrastructure at scale alongside Terraform-provisioned cloud resources.
Write custom automation and tooling in Go to reduce manual work and eliminate operational risk.
Turnkey is building the infrastructure for the autonomous economy, providing programmable guardrails that enable organizations to operate with autonomy and control. Founded by the team behind Coinbase Custody, it is a deeply technical, low-ego, high-agency team of experts in cryptography, security, and systems.
Lead end-to-end technical engagements: Partner directly with engineering teams to diagnose, unblock, and resolve complex infrastructure challenges.
Execute critical migrations: Develop reference implementations, tooling, and guidance to transition teams off deprecated systems seamlessly.
Accelerate platform adoption: Act as primary technical contact for new teams onboarding to Planet's core infrastructure.
Planet designs, builds, and operates the largest constellation of imaging satellites in history, delivering unprecedented dataset via a cloud-based platform for commercial, environmental, and humanitarian sectors. A global company with offices in the US, Europe, and Slovenia, Planet values a people-centric culture and community.
Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.
Own the design, development, and operation of infrastructure and build/release pipelines.
Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.
Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Migrate and modernize critical infrastructure from legacy Beanstalk to EKS, managing networking, IAM, and functional parity.
Design and maintain compute and networking components such as Private Link, load balancers, and Core API used by multiple teams.
Ensure high availability and resilience of shared compute infrastructure through on-call rotations and cross-team collaboration.
VTEX is a composable and complete commerce platform that empowers brands, distributors, and retailers with flexibility and comprehensive solutions. With over 1,300 employees across 16 countries, VTEX fosters a challenge-driven environment and collaborative culture.
Contribute to platform and harness engineering, including CI/CD and developer tooling.
Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
Participate in on-call rotation and incident management to ensure platform uptime.
Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Own infrastructure end to end, including Kubernetes clusters, Docker builds, and DigitalOcean services.
Move production deploys toward safer, more frequent releases and manage capacity, autoscaling, and load balancing.
Operate PostgreSQL, Redis, and NATS under real traffic, keep Cloudflare tight, and document systems for the whole team.
Colonist is building the biggest digital board game platform on the internet, where players have spent over 3,000 years playing 60 million games on their Catan alternative. The company is a fully remote, asynchronous team spread across multiple continents, focused on crafting polished digital board games.
Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.
Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.
Design and build self-service platform capabilities for engineering teams.
Develop and maintain global hybrid infrastructure across bare-metal, Linux, and Kubernetes.
Automate operational tasks and improve observability and reliability.
The partner company builds globally distributed infrastructure and platform capabilities. It is a small, highly autonomous team with a strong focus on reliability and developer experience.