Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Lead a high-impact infrastructure team, evolving internal platforms and CI/CD systems to support large-scale engineering operations.
Drive automation initiatives and AI-driven practices to reduce operational complexity and improve developer experience.
Define and execute strategies for scalable infrastructure, cloud environments, and platform engineering.
The partner company is a technology organization focused on building infrastructure platforms that enable engineering teams to deliver software faster. It is a remote-first company with a collaborative culture and a focus on innovation and scalability.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.
Collaborate with Development and Architecture teams to build complex and highly available cloud environments.
Provide Level 3 technical support for internal teams, customers, and partners.
Design, implement, and maintain a secure and scalable infrastructure platform.
Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with their FHIR-based data liberation platform. They were #19 on Deloitte's Technology Fast 50 Ranking for 2024 and foster a culture of respect, inclusion, and diversity.
Own and evolve Quansight's cloud infrastructure across AWS, Azure, and GCP.
Lead infrastructure engagements for clients from scoping through delivery.
Contribute to open-source projects and participate in upstream communities.
Quansight is rooted in the Python data science community and helps companies build sustainable solutions on open-source software. The team is a small, collaborative, fully distributed group of open-source maintainers and engineers.
Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.
Solid experience in cloud platform engineering, infrastructure, or platform operations.
Strong production experience with Azure, AWS, or GCP environments, including networking, IAM, security, and observability.
Experience designing observability systems across logs, metrics, traces, dashboards, and alerting workflows.
Cresteo is a nearshore tech services company focused on people-first, honest approaches. Led by experienced professionals, the team collaborates with global clients and values transparency, innovation, and profit-sharing.
Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
Build the human-AI development loop, creating tooling and automation for coding agent workflows.
Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.
Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.
Design, build, and optimize reliable infrastructure for healthcare technology.
Improve scalability, reliability, and performance across distributed systems.
Collaborate with engineers and data professionals to shape modern infrastructure practices.
This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.
Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
Collaborate with senior engineers to design scalable solutions and enhance developer experience.
Participate in incident response and retrospectives to drive continuous improvement.
Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.
Design and scale reliable AWS infrastructure across multiple accounts, ensuring security and compliance.
Collaborate with software engineers to advise on infrastructure best practices and optimize CI/CD pipelines.
Automate critical infrastructure updates and enhance observability with monitoring tools like Prometheus.
Constructor is an AI-first ecommerce search and discovery platform that helps shoppers find the right products and enables leading global e-commerce brands to drive revenue and conversion gains. They are a fully remote, diverse team offering unlimited vacation, training budgets, and the chance to work with smart colleagues.
Design, implement, and maintain highly available and scalable infrastructure solutions.
Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
Automate infrastructure deployment, configuration management, and operational workflows.
The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Design, deploy, and manage scalable cloud infrastructure across AWS and Azure, including EC2, S3, RDS, and Kubernetes.
Automate cloud resources using Terraform, CloudFormation, and PowerShell, while implementing security controls and monitoring.
Collaborate with development teams to integrate CI/CD pipelines and support containerized workloads with Docker and Kubernetes.
Mind Computing supports the Department of Veterans Affairs by providing cloud engineering and DevOps solutions. The company fosters a collaborative culture with a focus on security, reliability, and innovation, though the team size is not specified.
Design, deploy, and manage highly available and secure cloud infrastructure across AWS and Azure.
Automate infrastructure provisioning using Terraform, CloudFormation, and ARM templates while implementing CI/CD pipelines.
Collaborate with development teams to build cloud-native solutions and ensure operational excellence through monitoring and improvements.
Jobgether is a job matching platform that uses AI-powered technology to connect candidates with employers. They facilitate remote hiring processes and support large-scale technology initiatives with a focus on efficiency and objectivity.
Lead deployment and operation of product infrastructure in federal environments within AWS.
Build and maintain scalable, secure cloud-native platforms using Kubernetes, Terraform, and GitLab CI.
Improve development and deployment processes, create tooling for telemetry, and foster documentation culture.
Horizon3.ai is a fast-growing, remote cybersecurity company that helps organizations proactively find and fix exploitable attack vectors. We are a team of former special ops cyber operators and engineers committed to a culture of respect, collaboration, ownership, and results.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.