Provide advanced technical support with a focus on root-cause analysis, security, and employee experience.
The company is a global, technology-driven organization that redefines internal IT as an engineering discipline. It fosters a culture of innovation, automation, and continuous improvement with a distributed workforce.
Design, build, and operate scalable infrastructure platforms for large-scale AI model training and inference.
Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
Develop software solutions and automation tools using Python and systems programming languages like Go or C++.
Our partner builds and operates foundational technology powering advanced AI training and inference workloads at scale. They offer a collaborative culture focused on innovation, engineering excellence, and continuous learning.
Lead the design and rollout of new platforms to minimize incidents and enable customer-facing features.
Deploy updates and improvements for both internal and end customer use cases while collaborating with engineering and operations teams.
Participate in an on-call rotation evenly distributed across the team in a primary/secondary pattern.
Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. Founded in 2019, the company operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by leading venture capital firms.