Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Ensure reliability, performance, and scalability of Backcountry's multi-cloud platform.
Drive incident resolution, postmortems, and automation to reduce operational toil.
Leverage AI-assisted engineering tools and collaborate with teams to build and maintain observability and SLI/SLO instrumentation.
Backcountry is an online retailer of outdoor gear and apparel, rooted in adventure and the outdoor lifestyle. The company fosters a culture of recognition, wellbeing, and connection, with a lean, fast-paced engineering team.
Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Collaborate with cross-functional teams to design, implement, and maintain scalable and reliable infrastructure.
Utilize Infrastructure as Code (IaC) principles to automate provisioning, configuration, and deployment processes.
Troubleshoot and resolve complex technical issues related to infrastructure and deployment.
Pano AI is the leader in early wildfire detection and intelligence, helping fire professionals respond to fires faster and more safely using a combination of advanced hardware, software, and AI. The company is a 175+ person growth-stage hybrid-remote startup headquartered in San Francisco, recognized as one of the most innovative AI companies by Fast Company and TIME.
Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.
Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).
Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.
Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.
Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.
Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.
Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.
Build end-to-end internal tools across Python/Go services on GCP Cloud Run, TypeScript/React frontends, and Pub/Sub event-driven integrations.
Integrate with third-party SaaS platforms and cloud APIs, handling REST, webhooks, OAuth, and event-driven patterns with reliability.
Use AI tooling across the development lifecycle, directing Claude Code as a coding collaborator while owning the correctness, security, and output.
DoiT is a global technology company that helps organizations leverage the cloud for business growth and innovation, combining data, technology, and human expertise. With over 4,000 customers worldwide and decades of multi-cloud experience, the company fosters an entrepreneurial culture with remote flexibility.
Manage the lifecycle of internal Platform-as-a-Service and Data-as-a-Service products, overseeing highly technical engineering teams.
Define error budgets for 99.99% availability, lead cloud scaling strategies, and own the GCP cloud computing budget.
Partner with Data Engineers to optimize data pipelines and architecture, and ensure compliance with privacy regulations.
Sardine is the leading agentic risk platform for fighting financial crime, with an integrated solution unifying data across risk teams. The company has hubs in multiple locations but maintains a remote-first culture, hiring self-motivated individuals with extreme ownership and a high growth orientation.
Own observability end to end and define how we measure reliability.
Own CI/CD pipelines and make shipping fast and safe.
Footprint builds Percy, an AI agent that runs financial crime investigations end to end. The company is backed by QED, Index, and other investors, and its small, senior team ships fast and grew revenue 5x in the past year.
Collaborate with engineering teams to design scalable, secure systems.
Establish SLOs, manage incident response, and drive reliability improvements.
Leverage expertise in Go, Python, Kubernetes, and cloud platforms.
ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.