Source Job

US

  • Own 24×7 NTN RAN health, diagnose RF and protocol-level issues, and execute runbooks for fault recovery.
  • Serve as L3 escalation authority for RAN incidents, leading troubleshooting bridges and delivering root cause analysis.
  • Define RAN KPIs, optimize performance, and author operational runbooks and standards.

Kubernetes Python

20 jobs similar to Staff Network Reliability Engineer, RAN Operations

Jobs ranked by similarity.

US

  • Coordinate network outages with Voice, IP, Transport, and Local Operations teams to ensure quick resolution.
  • Perform event notifications for internal teams and external customers during outages.
  • Diagnose and resolve escalated network performance issues while supporting root cause analysis.

Lumos is transforming the future of connectivity by building one of the nation's fastest-growing 100% Fiber Optic networks. They are expanding access to high-speed internet with a fast-paced, collaborative culture focused on making an impact.

US

  • Perform telecommunications engineering studies and technical analyses to support the development and operation of a nationwide public safety broadband network.
  • Analyze wireless coverage, capacity, and performance data for 4G LTE and 5G networks using industry-standard tools and drive improvements.
  • Collaborate with government stakeholders, technical teams, and industry partners to validate solutions and enhance network reliability.

The company is a partner organization supporting the evolution and operation of a nationwide public safety broadband network. It focuses on mission-critical communications and fosters a collaborative environment driven by innovation and technical excellence.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

UK

  • Help define and mature Engineering Operations by improving application health visibility, service reliability, and operational analytics.
  • Build and implement scalable processes for Incident, Problem, and Change Management that engineers actually want to use.
  • Connect engineering systems, data, and teams to reduce fragmentation and improve operational visibility across the organization.

Turnitin is a recognized innovator in global education, developing learning integrity solutions that help educators and institutions uphold academic integrity. With over 16,000 academic institutions using our services in more than 185 countries, we foster a remote-first culture and a diverse community of colleagues across 35+ countries.

Global

  • Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
  • Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
  • Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.

Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.

Bulgaria

  • Design, implement, and support enterprise LAN, WAN, WLAN, SD-WAN, and cloud networking solutions.
  • Provide Tier 2/3 support for complex network incidents and outages, ensuring network availability and performance.
  • Lead technical aspects of network projects, mentor junior engineers, and collaborate with cybersecurity teams.

Sutherland is a global provider of business process and technology management services, serving clients across various industries. With over 60,000 employees worldwide, we foster a culture of innovation, collaboration, and continuous improvement.

  • Lead and mentor a high-performing field engineering team, defining deployment workflows and integration playbooks for repeatability and reliability.
  • Set technical strategy for field integrations, implementing scalable solutions and improving field engineering tooling with scripts and automation.
  • Partner across engineering, product, security, and mission operations to ensure secure, reliable deployments in customer-owned environments.

TurbineOne builds Mission-AI for the Frontlines, providing a Frontline Perception System that helps military and national security operators detect threats and accelerate decision-making at the tactical edge. The team is composed of experienced technologists, veterans, and operators committed to advancing national security through responsible innovation.

$85,000–$100,000/yr
Global

  • Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
  • Research network events to identify technical debt and drive improvements across platforms.
  • Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.

$197,000–$246,000/yr
United States Unlimited PTO 20w maternity 16w paternity

  • Set the vision and roadmap for the Tailscale dataplane, including WireGuard behavior, DERP relay expansion, and client-side networking.
  • Define performance and reliability as customer-facing product requirements, and own connectivity product strategy end to end.
  • Partner with GTM, Customer Success, and Solutions Engineering on production traffic commitments and strategic enterprise deals.

Tailscale builds software that securely interconnects people and devices, creating an easy-to-use, secure internet experience. Founded in 2019, the fully distributed company is backed by Accel, CRV, Insight, Heavybit, and Uncork Capital, and serves teams of all sizes.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

Global

  • Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
  • Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
  • Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.

US

  • Monitor network activity, respond to alarms and alerts, and manage tickets for broadband service providers.
  • Perform initial triage and basic troubleshooting of network issues related to routing, switching, and service availability.
  • Communicate clearly with customers, engineers, and vendors while tracking incidents through resolution and ensuring proper documentation.

Vantage Point Solutions is a customer-focused engineering and consulting firm that helps clients solve complex infrastructure challenges across broadband, power, and financial sectors. The company fosters a culture of collaboration, accountability, and professional growth.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

Canada Unlimited PTO 20w maternity 16w paternity

  • Own high-severity technical escalations from intake through resolution or engineering handoff.
  • Partner with Support and Product/Engineering to close the gap between customer problems and engineering fixes.
  • Monitor patterns across escalations to catch systemic issues and translate into product improvements.

Tailscale builds software that makes it easy to securely interconnect people and devices. Founded in 2019, the company is fully distributed and backed by Accel, CRV, and others.

  • Provide technical support for telecom transport issues, troubleshooting Layer 1 and Layer 2 OSI model and coordinating with carriers.
  • Manage inbound and outbound technical calls, create internal and carrier trouble tickets, and ensure resolution within SLAs.
  • Document interactions and escalate complex issues to senior technicians as necessary.

AireSpring is a leading provider of Cloud Communications, Managed Connectivity and Managed Security. The family-owned company has over 22,000 enterprise locations worldwide and is known for its award-winning customer service and integrity.

US Unlimited PTO

  • Serve as the first responder for production incidents, triaging and resolving issues.
  • Monitor application health and system availability using Datadog.
  • Develop automation scripts using Python or PowerShell to improve operational efficiency.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. Recognized as one of the fastest-growing companies in America, it offers a fulfilling work environment with career advancement opportunities across multiple locations in the US, South America, and India.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

United States

  • Diagnose and resolve complex network and voice routing issues using advanced troubleshooting and data analysis techniques.
  • Collaborate with cross-functional teams to improve processes and restore critical services in a 24/7 environment.
  • Mentor team members and document operational records to ensure long-term network stability and compliance.

This partner company specializes in network and telecommunications services, focusing on maintaining large-scale infrastructure. It operates with a collaborative and innovative culture, supporting career growth and inclusive workplace practices.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

UK

  • Lead cross-team incident triage for high-impact customer outages, coordinating Engineering, Product, and Customer Experience response and contributing to root cause analysis.
  • Develop and maintain observability for cloud-hosted customer deployments by building and refining system monitors, dashboards, and alerting.
  • Serve as the senior escalation point for complex support cases in EMEA, working cases that involve deep platform internals and unusual failure modes.

Dragos defends industrial organizations that provide modern civilization necessities like water, electricity, and safe working environments. As a market leader in ICS/OT Cybersecurity, we operate globally with a remote-first culture and are looking for mission-oriented teammates who value authenticity, transparency, and trust.