Collaborate with engineering teams to design scalable, secure systems.
Establish SLOs, manage incident response, and drive reliability improvements.
Leverage expertise in Go, Python, Kubernetes, and cloud platforms.
ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.
Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.
Build and operate large-scale cloud infrastructure for product metrics and real-time analytics.
Design, develop, and improve fault-tolerant distributed systems processing trillions of events.
Collaborate cross-functionally, perform code reviews, and participate in on-call rotation.
Jobgether uses AI-powered matching to connect candidates with hiring companies. They process applications quickly and share shortlists directly with employers, operating as a job platform seeking to streamline recruitment.
Own and operate 100+ multi-cloud streaming clusters and related database infrastructure in production.
Diagnose and eliminate cross-layer failure modes such as object storage latency, noisy neighbors, and query performance regressions.
Design safe upgrade and rollout strategies at scale, improving observability, automation, and operational ergonomics.
Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform built for scale. With over 35 million users and 7,000+ customers, we are a 100% remote company of 1,600+ team members across 40+ countries, backed by leading investors.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Co-own the architecture of cloud infrastructure on Azure and Kubernetes clusters for high throughput and availability.
Drive resilience strategy for global scaling, zero-downtime deployments, and disaster recovery.
Evolve observability stack with LGTM (Loki, Grafana, Tempo, Mimir) and lead incident response.
Flip is an AI-powered employee experience platform for frontline workers in retail, manufacturing, and logistics. The company is a young, rapidly growing tech company with a remote-first culture and offices in Berlin and Stuttgart.
Design and implement scalable sandboxed execution environments for intelligent agents and code execution systems.
Own the full development lifecycle from architecture and implementation to deployment and maintenance.
Collaborate closely with research, engineering, and platform teams to deliver solutions that enable experimentation and product development.
The company builds cutting-edge AI infrastructure, focusing on secure and scalable execution environments for intelligent systems. It is a collaborative and research-driven organization with a people-first culture that values diversity, curiosity, and collaboration.
Design, develop, and maintain scalable backend services in Go for data ingestion, storage, and querying.
Own and evolve critical platform components including APIs, messaging systems, and query-serving architectures.
Drive architectural decisions for system reliability, scalability, observability, and operational excellence.
Our partner is a company building foundational systems for large-scale data platforms and analytics capabilities, focusing on distributed backend services. They offer a remote-first, high-ownership environment with cross-functional collaboration and AI-driven productivity tools.
Own core compute infrastructure across multiple cloud providers and regions.
Design capabilities for greater performance and flexibility in service deployment.
Investigate and resolve challenging cloud and compute issues across the stack.
Render is a cloud platform for developers building AI-native, full-stack, multi-service applications. Trusted by over 6 million developers, the company has raised $257M in funding and values craft, velocity, and user experience.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Perform operational deployments, implementations, and maintenance for production systems.
Implement and maintain monitoring, reporting, and alerting systems for Core Speech products.
Be part of an on-call rotation and work collaboratively to improve system performance and architecture.
Solventum is a new healthcare company with a long legacy of solving big challenges to improve lives and enable healthcare professionals to perform at their best. They are a large company that values empathy, insight, and clinical intelligence, collaborating with top minds in healthcare.
Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.
Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.
Collaborate with Development and Architecture teams to build complex and highly available cloud environments.
Provide Level 3 technical support for internal teams, customers, and partners.
Design, implement, and maintain a secure and scalable infrastructure platform.
Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with their FHIR-based data liberation platform. They were #19 on Deloitte's Technology Fast 50 Ranking for 2024 and foster a culture of respect, inclusion, and diversity.
Support engineering teams by developing resilient applications on GKE and advising on best practices.
Develop and maintain infrastructure-as-code using Terraform, ArgoCD, and Python.
Manage GKE environments across multiple regions, including IAM and identity management.
Mimica uses AI-powered task mining to observe employee actions and create process maps, helping enterprises improve efficiency. The company is a fast-growing scale-up with a lean, collaborative culture.
Lead deployment and operation of product infrastructure in federal environments within AWS.
Build and maintain scalable, secure cloud-native platforms using Kubernetes, Terraform, and GitLab CI.
Improve development and deployment processes, create tooling for telemetry, and foster documentation culture.
Horizon3.ai is a fast-growing, remote cybersecurity company that helps organizations proactively find and fix exploitable attack vectors. We are a team of former special ops cyber operators and engineers committed to a culture of respect, collaboration, ownership, and results.
Design and scale highly reliable platform systems supporting complex cloud-native workloads across multiple deployment environments.
Build and enhance core platform services while contributing to distributed systems, event-driven architectures, and cloud-native infrastructure.
Optimize cloud resources, networking, storage, compute, and observability to improve system performance, scalability, reliability, and maintainability.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They operate as a job platform, processing applications and sharing top candidates with employers.
Run, upgrade, and evolve a multi-AZ ClickHouse estate on Kubernetes with rehearsed backup and recovery.
Design tables, sorting keys, materialised views, and optimize queries against multi-terabyte datasets.
Implement capacity models, storage tiering, and Grafana monitoring to keep the plant healthy.
Reactive Markets is a 2026 OTC Trading Platform of the Year handling over $50 billion in daily volumes across FX, Equities, and Crypto. Their engineering team is small, senior, and deeply invested in building reliable, high-performance systems.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.