Architect, build, and operate secure, multi-account AWS environments using modern Infrastructure as Code.
Design and optimize container orchestration with AWS ECS (Fargate) and Kubernetes (EKS) for specialized workloads.
Build unified CI/CD pipelines with GitHub Actions, embed security by design, and establish observability with Prometheus, Grafana, and CloudWatch.
Berlitz is a global language education company with a history of nearly 150 years, currently undergoing a digital transformation. The company fosters a remote-first, AI-native culture with a focus on autonomy and greenfield development, though employee count is not specified.
Lead the SRE strategy and execution for a high-growth AI company.
Build and scale a high-performing SRE team while defining reliability standards.
Architect secure, scalable cloud infrastructure and implement observability practices.
This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.
Design, maintain, and support secure AWS environments across compute, storage, networking, account structure, and operational practices.
Administer Azure-hosted applications, data services, and supporting infrastructure including tenant configuration and access controls.
Implement and maintain infrastructure as code using tools such as Terraform or CloudFormation to improve consistency and reliability.
JerseySTEM is a mission-driven professional network of pro-bono contributors dedicated to improving access to STEM education and career pathways for underserved middle school girls in New Jersey. Members contribute their professional skills and leverage their networks in service of the organization's gender-equity agenda, operating remotely with a small volunteer base.
Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.
Design, build, and operate multi-region AWS infrastructure on Kuberneties with Terraform and Helm at 15PB+ scale.
Own high-availablity, event-driven architectures and cost optimization across the stack.
Drive developer experience, security, and incident response as the second platform team member.
ScorePlay is the AI-powered media infrastructure for sports, automating content operations for the world's biggest sports organizations. We are a 50-person remote-first team based in New York and Paris, growing 2x year over year with 98% retention.
Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.
Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.
Design, implement, and manage secure cloud infrastructure for mission-critical applications.
Apply infrastructure-as-code practices using tools like Terraform or CloudFormation.
Collaborate with development and operations teams to optimize performance and reliability.
They provide cloud engineering and cybersecurity solutions for critical national security missions. They are a mission-focused, equal opportunity employer with a fully remote team.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Own infrastructure as code and build self-service paths for product teams.
Make reliability a property of the delivery path with SLOs and alerting.
Shift Left security and manage cloud costs as an engineering responsibility.
what3words is a global addressing company that assigns unique three-word addresses to every 3m square on Earth, making locations precise and easy to share. Their technology is used by emergency services, delivery companies, and automakers across 193 countries, with a growing user base and a microservices architecture on AWS.
Build and maintain reliable, scalable, and secure infrastructure solutions for large-scale SaaS applications.
Automate infrastructure provisioning, configuration, deployment, and optimization processes.
Collaborate with R&D teams to improve production stability, reliability, and developer experience.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use technology to ensure fair and efficient recruitment.