Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, build, and optimize internal platforms, automation workflows, and business system integrations.
Develop automation scripts, internal tools, and API-driven integrations using modern languages.
Implement security controls, identity management, and compliance processes across business systems.
This company is a globally distributed organization that builds and optimizes internal technology ecosystems. They have a collaborative international team and an engineering-driven, remote-first culture.
You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
You will foster a culture of quality and ownership by setting or improving code review and design standards.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Provide in-depth software engineering expertise in cloud architecture, design patterns, and programming.
Implement core DevOps practices, including Infrastructure as Code (IaC), continuous integration, and automated deployments.
Build and manage robust CI/CD workflows and drive architectural design and engineering best practices.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, it surpassed $3B in revenue in its last fiscal year and became a public benefit corporation in 2021, balancing interests of customers, employees, society, and investors.
Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.
PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.
Debug complex application systems to resolve business-impacting issues.
Lead system upgrades, migrations, and maintenance activities to ensure stability.
Apply DevOps practices and tools to deploy and run high-quality software.
Thoughtworks is a leading technology consultancy that helps clients solve complex business problems with technology. For over 30 years, it has fostered a dynamic and inclusive community of bright, supportive colleagues pushing boundaries through impactful work.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Alternative Payments is building the financial operating system for SMBs, consolidating the disconnected tech stack that holds service-based businesses back. We're growing fast, thinking big, and building a global team that wants to be part of something that lasts.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.