Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.
XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.
Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
Contribute upstream changes to open source projects and share on-call responsibilities.
Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.
ClickHouse provides a real-time analytics database for data warehousing, observability, and AI workloads. Recognized on the 2025 Forbes Cloud 100 list, it has over 4,000 customers and significant year-over-year growth.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Define and implement observability strategies, standards, and governance across applications and platforms.
Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
Collaborate across teams to improve reliability, deployment complexity, and incident response.
Contribute to roadmap planning, technical design, and long-term simplification of stack operations.
Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
Deliver engaging technical demos and presentations tailored to engineering and executive audiences.
Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.