Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.
Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.
Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.
Own and improve service reliability for the product team: design for HA/performance/scale, define SLIs/SLOs
Align with org standards, implement DevOps-driven updates: support processes, templates, services, breaking changes, security fixes
Build and evolve GitLab CI/CD for build, test, security scans, and progressive delivery; speed up and harden pipelines
Plata Card is a fintech company focused on cards and accounts services. They foster a high-tech environment with a supportive team and innovative spirit.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Manage large-scale, high-availability distributed systems on AWS with Terraform and CI/CD pipelines.
Implement robust monitoring and observability, practicing SRE and security compliance best practices.
Peek is the operating system powering the experiences industry, helping museums, attractions, and tours increase revenues and deliver seamless guest experiences. Recognized by Forbes as a Best Startup Employer and by Built In as a Best Place to Work, we are a global remote-first team of Peeksters who obsess over customers and collaborate with purpose.
Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.
PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.
Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.
Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.
Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Design and operate scalable cloud infrastructure across AWS and GCP.
Build and improve Kubernetes, Linux, and cloud networking environments.
Strengthen security, disaster recovery, and platform resilience.
Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.
You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
You will foster a culture of quality and ownership by setting or improving code review and design standards.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.
Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
Collaborate with senior engineers to design scalable solutions and enhance developer experience.
Participate in incident response and retrospectives to drive continuous improvement.
Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.