You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.
Zerohash is the leading crypto and stablecoin infrastructure platform, powering the next generation of financial services for banks, brokerages, fintechs, and payment companies. Founded in 2017, the company has raised over $280 million from top venture firms and strategic investors, and is trusted by global brands like Morgan Stanley and Stripe, operating with a compliance-first approach.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Own and evolve Quansight's cloud infrastructure across AWS, Azure, and GCP.
Lead infrastructure engagements for clients from scoping through delivery.
Contribute to open-source projects and participate in upstream communities.
Quansight is rooted in the Python data science community and helps companies build sustainable solutions on open-source software. The team is a small, collaborative, fully distributed group of open-source maintainers and engineers.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.
Design and operate scalable cloud infrastructure across AWS and GCP.
Build and improve Kubernetes, Linux, and cloud networking environments.
Strengthen security, disaster recovery, and platform resilience.
Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.
Design, build, and operate services and automations to manage Kubernetes clusters at scale, partnering with product management and technical leadership.
Drive rigorous code reviews and maintain high testing standards across the platform.
Manage cloud configurations across AWS and Azure using Terraform, ensuring deep observability and reliability.
Twilio is shaping the future of communications by delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers. They are a remote-first company with a strong culture of connection and global inclusion, employing a vibrant and diverse team.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.
ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.
Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.
Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.