Ensure reliability, scalability, and performance of cloud-based systems using Kubernetes and observability tools.
Define and monitor reliability metrics (SLIs, SLOs, MTTR) to continuously improve operational performance.
Automate operational tasks and implement Infrastructure as Code to reduce manual work and enhance efficiency.
Our partner is a technology company focused on building and maintaining reliable, scalable digital environments. They promote a culture of continuous improvement, collaboration, and proactive engineering.
Define SLIs, SLOs, and reliability targets for the platform.
Improve observability, alerting, and production readiness across services.
Automate operational work and support cloud/Kubernetes infrastructure.
Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.
Design and optimize AWS cloud infrastructures, managing Kubernetes clusters and implementing Infrastructure as Code with Terraform.
Administer observability stacks including Elasticsearch/OpenSearch, VictoriaMetrics, Grafana, Loki, and Vector ingestion pipelines.
Automate CI/CD pipelines using Python and Groovy, and resolve complex incidents through Jira and ServiceNow.
Devoteam is a European consulting firm specializing in digital strategy, technology platforms, cybersecurity, and business transformation. With over 10,000 employees across 20 countries in Europe, the Middle East, and Africa, the company combines enterprise-grade technology with a close-knit, professional team culture.
Design and own the foundational cloud and source control infrastructure for Webflow's corporate systems.
Build multi-account architecture, networking, DNS, PKI, and secrets management with self-service delivery pipelines.
Create guardrails through policy-as-code, cost controls, and security baselines to enable safe self-service for builders.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. With a remote-first culture and a focus on craft, Webflow is a growing company that values ownership, collaboration, and innovation.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Own infrastructure as code and build self-service paths for product teams.
Make reliability a property of the delivery path with SLOs and alerting.
Shift Left security and manage cloud costs as an engineering responsibility.
what3words is a global addressing company that assigns unique three-word addresses to every 3m square on Earth, making locations precise and easy to share. Their technology is used by emergency services, delivery companies, and automakers across 193 countries, with a growing user base and a microservices architecture on AWS.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Operate and optimize Amazon ECS deployments for JLV backend microservices.
Design and maintain CI/CD pipelines for containerized microservices with security compliance.
Implement observability and SRE practices to ensure high availability and incident response.
VetsEZ is a technology company that supports mission-critical clinical applications for the Department of Veterans Affairs, aggregating patient data from VA and DoD systems. As a growing government contractor, VetsEZ fosters a culture of innovation and compliance, providing remote work opportunities and professional development.
Own operational excellence for the Databricks platform, including monitoring, alerting, and incident response.
Design and maintain CI/CD standards for Databricks assets and production reliability patterns.
Partner with teams to ensure ingestion, observability, and compliance in a regulated environment.
Shield AI is a venture-backed defense-tech company protecting service members and civilians with intelligent systems. With offices across the U.S., Europe, the Middle East, and Asia-Pacific, its technology supports operations worldwide.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.