Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.
Design, build, and operate multi-region AWS infrastructure on Kuberneties with Terraform and Helm at 15PB+ scale.
Own high-availablity, event-driven architectures and cost optimization across the stack.
Drive developer experience, security, and incident response as the second platform team member.
ScorePlay is the AI-powered media infrastructure for sports, automating content operations for the world's biggest sports organizations. We are a 50-person remote-first team based in New York and Paris, growing 2x year over year with 98% retention.
Manage Kubernetes clusters using Rancher RKE2 and configure Calico CNI for networking.
Implement CI/CD pipelines with Jenkins, Terraform, and Ansible, integrating tools like Vault and Artifactory.
Ensure backup, recovery, and disaster recovery across hybrid cloud and on-prem infrastructure.
The company is hiring a Senior Kubernetes Engineer for an on-premises failover environment and critical application onboarding across clearing, settlement, and risk platforms. The team size and culture are not specified in the posting.
Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.
Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.
Support deployment, operation, and reliability of production services on Kubernetes.
Monitor service health, investigate production incidents, and participate in on-call and postmortems.
Troubleshoot application runtime, networking, and service-to-service issues across Node.js and JVM.
Software Mind develops innovative solutions for global companies, partnering with tech giants and unicorns on transformative projects. They foster cross-functional engineering teams with a culture of openness, respect, and passion, combining employment with enjoyment.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Own infrastructure as code and build self-service paths for product teams.
Make reliability a property of the delivery path with SLOs and alerting.
Shift Left security and manage cloud costs as an engineering responsibility.
what3words is a global addressing company that assigns unique three-word addresses to every 3m square on Earth, making locations precise and easy to share. Their technology is used by emergency services, delivery companies, and automakers across 193 countries, with a growing user base and a microservices architecture on AWS.
Design and evolve cloud infrastructure on GCP for scale and resilience.
Build internal tooling and automation that promote team autonomy and developer productivity.
Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.
The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.
Design and implement infrastructure using Terraform, Python, and Kubernetes on AWS.
Collaborate with engineering and data science teams to improve cloud infrastructure.
Automate CI/CD pipelines and enforce security governance and compliance.
Lyra Health is a mental health care provider serving 20 million people through employer and health plan partnerships. The company has delivered 15 million sessions and published 35 peer-reviewed studies, with a culture focused on clinical effectiveness.
Own infrastructure as code across development, staging, and production environments
Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
Manage cloud infrastructure, establish scalable engineering practices, and lead incident response
CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.