Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware using Python and infrastructure automation tools.
Engineer and support Python-based agents, APIs, and Ansible automation for hardware provisioning, telemetry, and health monitoring.
Analyze workload performance, thermals, and diagnostic output to identify hardware issues and improve validation methodologies.
Design, build, and maintain OS images for Vultr's Cloud Compute and Bare Metal platforms across Windows, Linux, and BSD distributions.
Convert legacy boot-up configuration scripts to cloud-init for consistent provisioning across all operating systems.
Manage the full lifecycle of marketplace applications, from onboarding new apps to updating and maintaining existing ones.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is a privately-held company with a $3.5 billion valuation, known for its flexible, scalable solutions and a culture of self-funding and growth.
Creating a golden ISO for imaging servers prior to shipping.
Performance testing Apache Traffic Server.
Building automation for server provisioning, patching, and monitoring/alerting.
ELEVI provides services to Federal and Commercial clients. They foster a culture of trust, empowerment, and diversity, offering flexible work arrangements and competitive benefits.
Lead and mentor a high-performing field engineering team, defining deployment workflows and integration playbooks for repeatability and reliability.
Set technical strategy for field integrations, implementing scalable solutions and improving field engineering tooling with scripts and automation.
Partner across engineering, product, security, and mission operations to ensure secure, reliable deployments in customer-owned environments.
TurbineOne builds Mission-AI for the Frontlines, providing a Frontline Perception System that helps military and national security operators detect threats and accelerate decision-making at the tactical edge. The team is composed of experienced technologists, veterans, and operators committed to advancing national security through responsible innovation.
Own major product components and engineering workstreams from discovery to production.
Design scalable architectures and build production-grade software with focus on reliability and security.
Use AI-assisted development tools to improve engineering efficiency.
They develop Linux security products and automation systems. They are a remote-first team emphasizing professional development and challenging technical projects.
Learn and contribute to production infrastructure automation using Ansible and Terraform.
Assist in building and maintaining cloud-native services across VKE, VLB, and VCR.
Develop hands-on skills with Kubernetes and container runtimes through guided project work.
Vultr provides high-performance cloud infrastructure solutions including Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage for enterprises and AI innovators worldwide. The company is the world's largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries, and values innovation and employee growth.
Operate and maintain on-premises infrastructure based on bare-metal Debian Linux servers.
Automate infrastructure provisioning and lifecycle management using Chef.
Manage and operate hybrid connectivity across GCP, AWS, Megaport, and on-prem data centers.
EXANTE is a global wealth-tech company powering the next generation of trading with centralized solutions and B2B financial infrastructure. The company brings together 600+ minds from 65 nationalities across 70 locations, fostering an informal, collaborative culture where trust and autonomy drive innovation.
Perform proactive maintenance and troubleshooting on data center server equipment.
Collaborate with suppliers, vendors, and customers to resolve hardware issues efficiently.
Maintain accurate documentation of hardware configurations and contribute to process improvements.
Gcore is a global provider of infrastructure and software solutions for AI, cloud, network, and security. With a team of over 550 professionals, they foster a collaborative culture focused on innovation and reliability.
Ensure smooth operation of the company's geo-distributed infrastructure.
Maintain and configure network equipment, bare metal servers, and NAS systems.
Administer multiple Linux servers and automate administrative tasks.
Xsolla is a global commerce company providing tools and services to help video game developers fund, distribute, market, and monetize their games. The company has helped over 1,500 game developers and fosters a community that values creativity, collaboration, and the transformative power of play.
Lead the design and rollout of new platforms to minimize incidents and enable customer-facing features.
Deploy updates and improvements for both internal and end customer use cases while collaborating with engineering and operations teams.
Participate in an on-call rotation evenly distributed across the team in a primary/secondary pattern.
Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. Founded in 2019, the company operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by leading venture capital firms.
Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
Research network events to identify technical debt and drive improvements across platforms.
Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Design, deploy, and manage RHEL-based environments across cloud and hybrid infrastructure.
Automate configurations using Ansible and develop scripts with Bash, Python, and PowerShell.
Ensure system security through hardening, patching, and vulnerability remediation.
Vattenfall is a European energy company that electrifies industries, supplies energy to homes, and modernizes living through innovation and cooperation. With approximately 21,000 employees, we foster a team-oriented culture focused on supporting a meaningful corporate mission.
Design and implement robust workload and network isolation for multi-tenant GPU bare-metal and virtualized environments.
Harden Linux kernel configurations, container runtimes, and orchestration layers against breakouts and privilege escalation.
Write low-level code in C, Go, or Rust to implement custom security controls, telemetry, and fixes at the OS and infrastructure level.
RunPod is the AI Developer Cloud, providing GPU cloud infrastructure for AI workloads. With over one million developers and a small, remote-first team, we prioritize ownership, speed, and impact, having closed a $100M Series A.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Deploy and support Scout products in real customer environments, building and improving deployment automations.
Prototype integrations with customer systems and turn repeatable issues into actionable product requirements.
Communicate clearly with technical and non-technical customers while researching cybersecurity trends.
Volexity is a cybersecurity company that builds products used in real-world security environments, including incident response and threat intelligence. The company values diversity and is an equal opportunity employer, hiring based on qualifications and merit.
Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.
Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.
Provide expert-level remote support for production and demo trading exchange environments to ensure system availability and stability.
Triage, troubleshoot, and drive resolution of production incidents, contributing to root cause analysis and post-incident reviews.
Assist with version upgrades, deployments, and release validation in coordination with Development and QA teams.
Crypto.com is the world's fastest growing global cryptocurrency platform, founded in 2016. It serves over 150 million customers and is committed to accelerating cryptocurrency adoption through innovation.
Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
Collaborate with compute and storage teams to support HPC and AI workloads.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Provide advanced technical support across hardware, software, and network issues, serving as an escalation point for complex troubleshooting.
Administer IT infrastructure including Okta, Google Workspace, and both macOS and Windows environments.
Lead process improvements, create documentation, and guide junior team members to enhance overall IT efficiency.
Openly redefines the insurance experience by replacing complexity with clarity, offering tailored coverage for modern homeowners through independent agents. They are a remote-first company that values collaboration, communication, and work-life balance, bringing together diverse problem-solvers to challenge the status quo.