Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Design and maintain CI/CD and MLOps pipelines for software and machine learning models.
Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters.
Champion Infrastructure as Code and observability to ensure high availability and governance.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive solutions and cloud capabilities. Headquartered in Singapore, the company has deployed data centers across multiple countries and fosters a culture of innovation.
Design and maintain CI/CD and MLOps pipelines for AI and software applications, ensuring seamless deployment and automation.
Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters to support high-performance AI workloads.
Champion Infrastructure as Code and observability practices to ensure high availability, security, and compliance across multi-cloud environments.
Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company has a global presence with data centers in multiple countries and a culture focused on innovation and reliability.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
Support development teams by consulting on infrastructure needs and improving observability and incident response.
Provide operational support and maintain platform reliability through structured debugging and on-call rotations.
PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.
Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
Own platform reliability, observability, capacity planning, and operational readiness.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.
Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Manage physical "Metal" environments from bare metal to Kubernetes, including cluster networking and scheduling.
Maintain Crossplane compositions and Terraform modules for cloud service provider resources.
Work with application teams to understand needs and invest in right capabilities.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. They are a 100% remote company with team members across 40+ countries, backed by leading investors, and known for an open-source legacy and global collaborative culture.
Define and lead platform engineering strategy across complex, multi-environment cloud systems.
Architect scalable Kubernetes platforms, own IaC standards, and drive DevSecOps implementation.
Mentor engineers, partner with leadership on infrastructure direction, and lead complex migrations.
Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, with delivery centers in Canada, the US, Eastern Europe, and Latin America, we are a nimble team of senior engineers averaging 15+ years of experience.
Manage and troubleshoot complex distributed large-scale software systems
Build scalable, secure and reliable container-based infrastructure
Automate software delivery processes with CI/CD pipelines
Coinspaid Dev is the engineering brand behind the technology, infrastructure, and R&D expertise built within Coinspaid, focusing on advancing blockchain infrastructure engineering. With over 120 engineers and more than 11 years of industry experience, they bring together teams building distributed systems and blockchain infrastructure across 20+ blockchain networks.
Partner with customers to plan and deliver migrations from self-managed GitLab to GitLab Dedicated, using GitLab Geo for data replication with minimal disruption.
Balance hands-on technical delivery with clear communication and project ownership, translating customer goals into practical plans for cutover and validation.
Improve tooling, scripts, and runbooks to make migrations repeatable and reduce manual work, sharing insights with Product and Engineering teams.
GitLab is an intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. More than 50 million registered users and over 50% of the Fortune 100 trust GitLab to ship better software, and the company fosters a high-performance culture driven by values and continuous knowledge exchange.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Own cloud infrastructure and Kubernetes environment, keeping it reliable, secure, and cost efficient.
Lead the team in using AI-assisted engineering to design, build, and operate platform infrastructure.
Manage and mentor a team of engineers while staying hands-on to contribute directly to the work.
Doma Technology provides solutions for lenders, real estate professionals, title agents, and homeowners that make closings simpler and more efficient. The company values an entrepreneurial, people-first culture with a focus on diversity, equity, and inclusion.
Design and optimize AWS cloud infrastructures, managing Kubernetes clusters and implementing Infrastructure as Code with Terraform.
Administer observability stacks including Elasticsearch/OpenSearch, VictoriaMetrics, Grafana, Loki, and Vector ingestion pipelines.
Automate CI/CD pipelines using Python and Groovy, and resolve complex incidents through Jira and ServiceNow.
Devoteam is a European consulting firm specializing in digital strategy, technology platforms, cybersecurity, and business transformation. With over 10,000 employees across 20 countries in Europe, the Middle East, and Africa, the company combines enterprise-grade technology with a close-knit, professional team culture.
Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.
Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Architect, deploy, and manage highly available, fault-tolerant cloud infrastructure across Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
Maintain and scale declarative infrastructure using Terraform across a multi-hundred-file estate, enforcing GitOps workflows with Atlantis.
Build, maintain, and optimize robust automated pipelines for continuous integration and delivery using GitHub Actions, Jenkins, and ArgoCD.
Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is a scrappy, nimble organization dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.