Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.
WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Design, build, and deploy production systems with focus on scalability, reliability, and security.
Develop and maintain automation to streamline operations and eliminate toil.
Proactively monitor systems and implement automated incident response to minimize downtime.
Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.
Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
Build observability platforms to detect issues proactively and mentor engineers on reliability standards.
Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.
Build and harden secure CI/CD pipelines with security gates to catch issues before production.
Lead security architecture reviews and threat models for Kubernetes-based workloads on GCP and AWS.
Harden container images, Kubernetes configurations, and cloud IAM to minimize attack surface.
Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. The company is venture-backed by leading investors and serves Fortune 500 enterprises, with a culture that values customer obsession, intentional action, and trust.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Design and deliver solutions for cloud-hosted production infrastructure, including CI/CD pipelines and automation.
Shape development and deployment of mission-critical enterprise software using optimized and automated CI/CD pipelines.
Design, build, and support cloud infrastructure and security technologies for resiliency, observability, and cost optimization.
Ping Identity provides an intelligent cloud identity platform that enables secure and seamless digital experiences for users. The company serves more than half of the Fortune 100, has offices globally, and fosters a culture of digital freedom and respect for individuality.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Lead CI/CD pipeline optimization and container orchestration to scale infrastructure across teams.
Architect secure infrastructure with secrets management, IAM, and vulnerability scanning as default.
Drive platform reliability through SLOs, observability, and incident response.
Evolve aims to make vacation rental easy for everyone. The team is high-performing, customer-obsessed, and runs on curiosity, communication, and accountability.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Turn developer, agent, and security problems into clear requirements, staged technical decisions, and measurable outcomes.
Design, build, and operate Go services and APIs that power Docker's build and secure-CI infrastructure.
Develop isolation, identity, policy, provenance, attestation, and signing capabilities for trusted build and test jobs.
Docker is a leading developer tooling platform trusted by over 20 million monthly users and 20 billion container image pulls. They are a globally distributed, remote-first team building tools that define how software gets built and delivered.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Write, tune, and maintain detections in a modern SIEM across cloud, container, and SaaS log sources.
Run cloud security operations in AWS, triage alerts, and help execute incident-response playbooks.
Build and extend security-automation tooling with code fluency in Python or TypeScript.
SmarterDx uses clinical AI to help health systems capture the full value of patient care. They are a remote-first team with a mission to make healthcare more accurate and sustainable.
Own the technical strategy for multi-ecosystem scaling, defining architecture for onboarding new language ecosystems.
Drive end-to-end remediation automation, leading redesign of CVE workflows to close the loop from detection to verified release.
Set platform-wide technical direction spanning package index, build pipelines, and orchestration tooling to serve customers and ecosystem teams.
Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. They serve Fortune 500 enterprises and global industry leaders, and are venture-backed by leading investors, fostering a culture of customer obsession and intentional action.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.