Design, build, and deploy production systems with focus on scalability, reliability, and security.
Develop and maintain automation to streamline operations and eliminate toil.
Proactively monitor systems and implement automated incident response to minimize downtime.
Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.
Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
Research network events to identify technical debt and drive improvements across platforms.
Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.
Design and implement enterprise-scale Network-as-Code and GitOps solutions using Terraform, Ansible, and Python.
Build and optimize CI/CD pipelines for network infrastructure with automated testing and zero-downtime deployments.
Lead design and troubleshooting of large-scale network architectures including BGP, OSPF, and Kubernetes networking.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1000 full-time professionals and a 99% project success rate, the company operates in over 25 countries with a culture of Relentless Performance.
Build and maintain automation platforms for large-scale hardware infrastructure operations.
Develop tools to eliminate manual processes and improve system reliability and visibility.
Collaborate with cross-functional teams to design, deploy, and optimize critical infrastructure services.
The company specializes in building and operating large-scale hardware automation platforms for AI infrastructure. They foster a collaborative, international team culture with a focus on technical growth and autonomy.
Design, implement, and maintain highly available and scalable infrastructure solutions.
Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
Automate infrastructure deployment, configuration management, and operational workflows.
The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.
Design and implement core networking services for Docker's Sandboxes platform across local and public cloud environments.
Build scalable networking infrastructure for microVM orchestration, workload scheduling, and lifecycle management.
Develop high-performance networking components and ensure reliability, observability, and performance across the infrastructure.
Docker is a developer tooling platform trusted by over 20 million monthly users and 20 billion container image pulls, enabling developers to build, share, and run applications. The company is a globally distributed, remote-first team with offices in Seattle and Paris, focused on innovation in AI and security.
Lead design, implementation, and troubleshooting of enterprise network infrastructure including routing, switching, firewalls, VPN, and wireless.
Own and mature observability strategy using Datadog, Dynatrace, LogicMonitor, and SNMP-based polling to build proactive monitoring and alerting.
Drive infrastructure as code with Terraform and GitHub-based CI/CD pipelines for scalable, secure, and repeatable IT operations.
EverOps is a premier Embedded Service Provider that partners with customer IT teams to address mission-critical delivery and infrastructure challenges. The company operates a U.S.-Based Virtual Operating Center and offers equity, unlimited PTO, and sponsored healthcare.
Ensure smooth operation of the company's geo-distributed infrastructure.
Maintain and configure network equipment, bare metal servers, and NAS systems.
Administer multiple Linux servers and automate administrative tasks.
Xsolla is a global commerce company providing tools and services to help video game developers fund, distribute, market, and monetize their games. The company has helped over 1,500 game developers and fosters a community that values creativity, collaboration, and the transformative power of play.
Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
Collaborate with compute and storage teams to support HPC and AI workloads.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.
Build and own the end-to-end zero-touch provisioning (ZTP) and automation platform for GPU network fabrics.
Lead a small team of engineers and SREs, ensuring clean integration with broader platform tooling.
Bring software engineering rigor to network automation including code review, testing, CI/CD, and release management.
TensorWave delivers seamless, secure, reliable, and resilient AI compute at scale by building a versatile cloud platform that eliminates infrastructure barriers. They empower builders to focus on innovation instead of fighting their stack, and are committed to creating an inclusive environment for all employees.
Lead complex cross-functional technical programs involving cloud networking and infrastructure services.
Translate high-level objectives into structured execution plans with milestones and timelines.
Collaborate with engineering managers and technical leads to align priorities and ensure smooth execution.
The company is a provider of cutting-edge cloud networking infrastructure for AI workloads. It fosters a collaborative, engineering-driven culture with a focus on scalability and resilience.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Manage end-to-end connectivity infrastructure projects from site evaluation to operational readiness.
Act as technical owner for fiber pathway architecture and coordinate with internal teams and external partners.
Build strategic relationships with telecom providers and support procurement activities.
The company designs and delivers connectivity infrastructure for cloud and AI platforms. The team is collaborative and innovative, with a focus on technical excellence and continuous improvement.
Design, develop, and support network communication solutions for secure and scalable environments.
Work with protocols like TCP/IP, HTTP, DNS, and cloud platforms such as AWS, Azure, and GCP.
Collaborate with engineering teams to optimize performance and reliability in distributed systems.
The partner company develops modern communication systems and protocol-driven solutions for secure, scalable network environments. They foster a technically driven, flexible remote work culture with a focus on engineering excellence.
Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.
Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Manage and resolve the most challenging issues for the SRE team, focusing on instance performance, reliability, and availability.
Use software development and systems engineering experience to proactively prevent issues and drive improvements in infrastructure reliability.
Drive a culture of automation and scalable solutions, collaborating with partner teams to enhance system design.
ServiceNow provides an AI control tower for business reinvention, integrating any AI, data, and workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture, combining technology and talent to drive innovation and growth.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Design, deploy, and operate enterprise-grade network infrastructure across global data centers and cloud environments.
Manage cloud networking, BGP routing, VPNs, F5 load balancing, and network automation with Terraform and Ansible.
Participate in on-call rotation to resolve critical network incidents and support hybrid cloud connectivity.
Applied Systems is an insurtech company that delivers innovative software and services to insurance agencies. With over 40 years of experience, they foster a culture built on values that make their team indispensable to each other.