Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Design and evolve cloud infrastructure on GCP for scale and resilience.
Build internal tooling and automation that promote team autonomy and developer productivity.
Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.
The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.
Serve as the first responder for production incidents, triaging and resolving issues.
Monitor application health and system availability using Datadog.
Develop automation scripts using Python or PowerShell to improve operational efficiency.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. Recognized as one of the fastest-growing companies in America, it offers a fulfilling work environment with career advancement opportunities across multiple locations in the US, South America, and India.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Alternative Payments is building the financial operating system for SMBs, consolidating the disconnected tech stack that holds service-based businesses back. We're growing fast, thinking big, and building a global team that wants to be part of something that lasts.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.
Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.