Source Job

Global 6w PTO 26w maternity 26w paternity

  • Design and write high-performing scalable software for training models.
  • Develop new tools to support and accelerate research and LLM training.
  • Collaborate with engineering teams and scientific teams to implement experiments on cluster and data infrastructure.

Python Distributed Training Kubernetes Machine Learning

20 jobs similar to Member of Technical Staff, Agentic Environments

Jobs ranked by similarity.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

Global 5w PTO

  • Contribute to the entire development cycle of cutting-edge large deep learning models, from dataset preparation to deployment.
  • Collaborate with engineers and researchers to translate research into practical applications, wearing many hats.
  • Work at an exciting moment to build something from the ground up with a kind and collaborative team.

Reka builds useful multimodal AI to empower organizations, as a globally distributed foundation model startup headquartered in San Francisco Bay Area. The team, including contributors from Google DeepMind and FAIR, embraces a remote-first culture and collaborates globally.

$100,000–$300,000/yr
US

  • Build the software backbone for autonomous foundation models, including multimodal data pipelines and training workflows.
  • Implement and iterate on LLM, VLM, and VLA architectures, owning model code paths and inference runners.
  • Deliver production-grade serving tooling for low-latency operation and own systems from architecture to iteration.

We're building the next generation of ground transportation with advanced physical AI to simplify freight challenges. Our stealth team, founded by engineers who scaled autonomous driving, is developing a new vehicle platform and focuses on creating reliable real-time autonomous systems.

Global

  • Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
  • Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Apply quantization, batching, caching, and routing to optimize latency and cost at scale.

vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.

US UK Unlimited PTO

  • Design and optimize training and post-training pipelines for large language models.
  • Improve model quality through supervised fine-tuning, reinforcement learning, and evaluation.
  • Build PyTorch-based training infrastructure and optimize distributed training across multi-GPU environments.

Lightning AI builds an end-to-end platform for developing, training, and deploying AI systems, founded in 2019. They are a global company with offices in New York, San Francisco, Seattle, and London, backed by major venture capital firms, and foster a builder culture that values urgency, ownership, and open communication.

Global

  • Taking ML or LLM proof-of-concept to production for large enterprises.
  • Designing and hardening data and training pipelines for enterprise ML systems.
  • Building LLM and RAG systems with retrieval quality, evaluation and cost control.

Janea Systems (USA) is a dynamic team of the best & brightest software engineering specialists and solutions innovators from around the world. From kernel to cloud, we provide high-impact software development services to Fortune 500 companies.

$216,700–$303,400/yr
US

  • Design, develop, and deploy ML models, including large language models, for various NLP tasks.
  • Collaborate with cross-functional teams to gather requirements, define architectures, and iterate on model development.
  • Stay up-to-date with latest research and contribute to best practices for responsible ML development.

Reddit is a community of communities built on shared interests, passion, and trust, hosting the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 126 million daily active users, it is one of the internet's largest sources of information, fostering a culture of authenticity and community.

  • Build and ship AI features end-to-end, from model to system to user experience.
  • Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
  • Debug full-stack issues and optimize for latency, cost, and production performance.

A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.

Global 6w PTO 26w maternity 26w paternity

  • Conduct cutting-edge machine learning research, building and training large language models.
  • Focus on research projects aimed at expanding the frontier of knowledge in language modelling and associated areas.
  • Disseminate your research results through publications, datasets, and code.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprise AI systems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, and more.

$153,351–$206,481/yr
Canada

  • Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
  • Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
  • Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.

Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

$200,000–$350,000/yr
US

  • Develop and deploy machine learning and AI systems.
  • Work with LLMs, generative AI, and modern ML frameworks.
  • Optimize model performance, latency, and cost.

A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.

France

  • Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
  • Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
  • Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.

Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

Global

  • Design and scale production ML systems for LLM-based applications.
  • Build training and evaluation pipelines for continuous model improvement.
  • Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.

A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.

Global 6w PTO 26w maternity 26w paternity

  • Manage the program portfolio covering inference, efficiency, serving, and endpoints to scale Cohere's infrastructure.
  • Lead cross-functional coordination with Modeling and customer-facing teams for end-to-end execution.
  • Identify pain points, establish processes, and improve engineering best practices.

Cohere is the leading security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are a global team of researchers, engineers, and designers passionate about our craft, with offices across North America and Europe.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Global 6w PTO 26w maternity 26w paternity

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
  • Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
  • Optimize performance without compromising model quality.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.

Canada 6w PTO 26w maternity 26w paternity

  • Develop and deliver cutting-edge agentic AI solutions using Cohere's foundation models and Agentic AI Foundry.
  • Architect scalable, secure, and customizable NLP and generative AI solutions for enterprise customers.
  • Collaborate with customers to understand complex workflows, design pilots, and translate business requirements into technical solutions.

Cohere is the leading security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are a global team of researchers, engineers, and designers passionate about our craft, headquartered in Toronto with offices worldwide.

$165,000–$330,000/yr
US Unlimited PTO

  • Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
  • Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
  • Work across product, software development, performance engineering, and customer-facing implementations.

Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.