Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.
Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.
Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
You will respond to support escalations and troubleshoot complex technical problems.
You will develop software, tools, and scripts to automate deployment and monitoring of production systems.
We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.
Zerohash is the leading crypto and stablecoin infrastructure platform, powering the next generation of financial services for banks, brokerages, fintechs, and payment companies. Founded in 2017, the company has raised over $280 million from top venture firms and strategic investors, and is trusted by global brands like Morgan Stanley and Stripe, operating with a compliance-first approach.
Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
Build observability platforms to detect issues proactively and mentor engineers on reliability standards.
Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.
Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
Monitor production environments to ensure high availability, reliability, and performance.
Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.
Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.
Develop and maintain infrastructure automation solutions using Ansible.
Design, implement, and enhance CI/CD pipelines and operational tooling.
Troubleshoot complex Linux-based infrastructure and distributed systems issues to maintain high availability.
itD is a consulting and software development company that blends diversity, innovation, and integrity with real business results. They are a woman- and minority-led firm offering a dynamic culture of respect, empowerment, and recognition.
Design, build, and optimize reliable infrastructure for healthcare technology.
Improve scalability, reliability, and performance across distributed systems.
Collaborate with engineers and data professionals to shape modern infrastructure practices.
This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.
Manage and troubleshoot complex distributed large-scale software systems
Build scalable, secure and reliable container-based infrastructure
Automate software delivery processes with CI/CD pipelines
Coinspaid Dev is the engineering brand behind the technology, infrastructure, and R&D expertise built within Coinspaid, focusing on advancing blockchain infrastructure engineering. With over 120 engineers and more than 11 years of industry experience, they bring together teams building distributed systems and blockchain infrastructure across 20+ blockchain networks.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Serve as technical lead for the Brain Team, engineering and operating enterprise Kubernetes platforms on Amazon EKS.
Design and optimize Kubernetes networking using Cilium and Istio, including service mesh and network policies.
Automate cloud infrastructure with Terraform and implement GitOps workflows for secure CI/CD pipelines.
VetsEZ is a technology company specializing in federal government healthcare modernization projects. They are an equal opportunity employer with a focus on building secure, automated platform capabilities for large-scale clients.
Lead the SRE strategy and execution for a high-growth AI company.
Build and scale a high-performing SRE team while defining reliability standards.
Architect secure, scalable cloud infrastructure and implement observability practices.
This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.