Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Serve as the first responder for production incidents, triaging and resolving issues.
Monitor application health and system availability using Datadog.
Develop automation scripts using Python or PowerShell to improve operational efficiency.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. Recognized as one of the fastest-growing companies in America, it offers a fulfilling work environment with career advancement opportunities across multiple locations in the US, South America, and India.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Design, implement, and maintain highly available and scalable infrastructure solutions.
Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
Automate infrastructure deployment, configuration management, and operational workflows.
The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.
SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Design and build cloud-native engineering platforms for software validation, release validation, and production readiness.
Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left practices.
Foster a culture of reliability, automation, and operational excellence while mentoring engineers.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help organizations work smarter. They serve 85% of the Fortune 500 and foster an AI-native culture where technology and talent are unstoppable.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.
Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.
Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.
Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.
Monitor, triage, and resolve customer-reported incidents within defined SLAs.
Serve as primary point of contact for technical issues related to installation, Helm configuration, and integrations.
Troubleshoot Kubernetes-related issues such as ingress, SSO, and other integrations.
ScaleOps redefines autonomous cloud and AI infrastructure, freeing DevOps engineers from manual resource management. The company is backed by over $210M in funding and trusted by leading enterprises including Adobe and Coinbase.
Drive deployment automation across EKS clusters using GitOps.
Own infrastructure requirements and coordinate maintenance backlog.
Collaborate with engineering teams and integrate AI agents to accelerate workflows.
Ping Identity provides an intelligent cloud identity platform that secures digital experiences. With global offices and serving over half of the Fortune 100, the company fosters a culture of respect and individuality.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.