Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.
Own and improve production infrastructure reliability and stability.
Prepare, execute, and support deployments and infrastructure changes.
Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Our international team of digital nomads works remotely from all over the world and we are proud to be a two-time 'Great Place to Work' winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Design, deploy, and manage scalable cloud infrastructure across AWS and Azure, including EC2, S3, RDS, and Kubernetes.
Automate cloud resources using Terraform, CloudFormation, and PowerShell, while implementing security controls and monitoring.
Collaborate with development teams to integrate CI/CD pipelines and support containerized workloads with Docker and Kubernetes.
Mind Computing supports the Department of Veterans Affairs by providing cloud engineering and DevOps solutions. The company fosters a collaborative culture with a focus on security, reliability, and innovation, though the team size is not specified.
Operate and evolve AWS infrastructure for Data/AI platforms, ensuring security, scalability, and high availability.
Build CI/CD pipelines, automate provisioning with Terraform, and implement observability.
Collaborate with Data, AI, and Infrastructure teams, document standards, and drive platform improvements.
The partner company is building a modern Data Platform team focused on secure, scalable, and highly available infrastructure for Data and AI workloads. The culture emphasizes DevOps, automation, and continuous improvement, with close collaboration across Data, AI, and Infrastructure teams.
Design and implement infrastructure using Terraform, Python, and Kubernetes on AWS.
Collaborate with engineering and data science teams to improve cloud infrastructure.
Automate CI/CD pipelines and enforce security governance and compliance.
Lyra Health is a mental health care provider serving 20 million people through employer and health plan partnerships. The company has delivered 15 million sessions and published 35 peer-reviewed studies, with a culture focused on clinical effectiveness.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.
Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.
Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.
Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.
Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.
Implement and maintain CI/CD pipelines, containerization, and Infrastructure as Code on AWS.
Develop monitoring, logging, and performance tools for microservices and APIs.
Collaborate with teams to resolve system issues and ensure security, scalability, and availability.
Bellese is a mission-driven digital services company focused on innovative technology solutions in civic healthcare to improve public health outcomes. They foster a collaborative, learning environment and are remote-first with a team-oriented culture.
Define and monitor reliability metrics such as SLI, SLO, SLA, MTTR, and MTTD.
Implement observability solutions including monitoring, alerting, dashboards, and APM.
Collaborate with multidisciplinary teams to embed reliability and observability into solutions.
The company is a technology organization focused on building and maintaining reliable digital environments. It fosters a culture of engineering excellence, collaboration, and data-driven decision-making.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.
Design, implement, and maintain secure CI/CD pipelines on AWS for a SaaS platform.
Implement infrastructure-as-code with security best practices and manage IAM.
Support SOC 2 compliance and document existing systems for knowledge transfer.
GoFasti is a Talent-as-a-Service company connecting world-class developers and designers from Latin America with global companies. They focus on remote work and building strong relationships with their talent pool.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Designs, engineers, and automates secure enterprise cloud environments supporting mission-critical government workloads.
Provides technical leadership across AWS GovCloud, cloud architecture, infrastructure automation, container platforms, and CI/CD.
Develops standardized hosting models and integrates AWS IaaS, PaaS, and SaaS capabilities with enterprise applications.
True Zero Technologies is a veteran-owned small business that enables people and technology to drive quality outcomes. With a people-first culture, it has been recognized as a Best Place to Work in 2023 and 2025, and made the Inc. 5000 list of fastest-growing companies in America in 2022, 2023, and 2025.
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.
YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.
Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.