Own the design, resilience, and health of infrastructure across AWS and hybrid on-prem/Azure/GCP environments.
Execute disaster recovery, patching, and security/compliance controls across a large, multi-account, multi-platform estate.
Build CI/CD pipelines, automate to reduce toil, and serve as L2 escalation for infrastructure incidents.
HealthEdge offers AI-powered operational infrastructure for health insurance companies, guaranteeing an enduring financial edge in an increasingly competitive market. We're experiencing strong market momentum with a growing number of health plans choosing HealthEdge to modernize their operations and competing effectively, making this a pivotal moment to join and shape the future of healthcare technology.
Design and operate scalable AWS infrastructure with containerization and orchestration tools.
Implement monitoring, logging, and infrastructure as code using Terraform.
Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.
Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.
Automate infrastructure creation using Terraform and AWS CloudFormation.
Own build and release cycles, and improve CI/CD pipelines with Jenkins and Bitbucket.
Migrate legacy infrastructure to AWS and build scalable data platforms like EMR and Redshift.
Mactores is an AWS modernization firm that uses its Aedeon agent platform to automate repetitive discovery and validation, enabling engineers to focus on architecture and cutover. The company operates with a forward-deployed model and a culture built on 10 core leadership principles like curiosity and ownership.
Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
Implement observability solutions and enforce security best practices to ensure uptime and system performance.
Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.
Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.
Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
Build and maintain automation for lifecycle management of Kubernetes clusters.
Identify and root-fix reliability bottlenecks before they become incidents.
Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Design, build, and optimize multi-region, high-availability AWS infrastructure.
Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
Build end-to-end telemetry and own incident management to harden reliability.
VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.
Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.
Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Own the design, development, and operation of infrastructure and build/release pipelines.
Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.
Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.
Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).
DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.
Own the internal infrastructure and support queue, from Google Workspace administration to cloud provisioning and Kubernetes deployments.
Support customer project deployments by executing runbooks, troubleshooting issues, and contributing to hardening and compliance work.
Collaborate with a fully remote, distributed team using asynchronous communication, with opportunities to grow into environment builds and platform ownership.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. The company is built by people with deep roots in open-source, with a culture of ownership, remote collaboration, and diversity.
Establish and employ continuous integration and delivery (CI/CD) patterns for successful software solutions.
Design secure, operationally sound solutions across AWS, Azure, OpenShift, and IBM Cloud.
Implement observability stacks, manage Kubernetes clusters, and automate infrastructure with Terraform.
Conga unifies commercial operations by aligning pricing, quoting, contracting, rebates, and communications so companies run as connected, smarter businesses. With more than 10,000 customers worldwide, including over 50% of the Fortune 100, Conga fosters a collaborative culture where every voice is heard.
Build and maintain scalable cloud infrastructure for high availability.
Enhance observability and monitoring frameworks for accurate alerts.
Support on-call rotations and incident response with post-mortems.
GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.