Source Job

  • Own the architectural evolution of the operator foundation model, designing transformer variants for spatial domains and scaling distributed training beyond 45TB datasets.
  • Architect trillion-voxel inference systems that balance memory, compute, and communication with production-grade stability.
  • Ship expanded operator capabilities into production, increasing simulation throughput by 100x across global, multi-entity deployments.

PyTorch Machine Learning Systems Engineering

20 jobs similar to Member of Technical Staff - Foundation Model Architecture & AI Infrastructure

Jobs ranked by similarity.

US

  • Design high-performance training platform components including orchestration, observability, and performance tuning.
  • Deliver end-to-end ML pipelines from dataset curation to training, validation, and deployment.
  • Drive technical direction across ML Platform, Infrastructure, Autonomy, and Safety teams.

Stack is developing revolutionary AI and advanced autonomous systems for safer, more reliable, and efficient operations, focusing on autonomous trucking. The Stack team brings decades of experience in real-world systems and is committed to a culture of inclusion, entrepreneurship, and innovation.

China

  • Operate and expand Telnyx's own B300 GPU fleet to maximize inference throughput per GPU-dollar.
  • Design and implement serverless inference for open-weight models and dedicated enterprise deployments.
  • Work upstream in open-source technologies like vLLM, SGLang, and Kubernetes.

Telnyx is an industry leader building the future of global connectivity through a private, multi-cloud IP network and edge APIs. The company is financially stable and profitable, with a global team and a focus on innovation and continuous learning.

$230,000–$330,000/yr
US Unlimited PTO

  • Own Merlin's foundation and world-model work, including architecture selection, post-training, and capability roadmap.
  • Lead and mentor a small team of world-model engineers, setting the technical bar and review culture.
  • Design model interface to the autonomy stack with structured, schema-constrained plan outputs and build evaluation harnesses.

Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it is expanding its organization to accelerate the deployment of its autonomy platform.

AI Engineer

Cyvl
US

  • Build and improve computer vision and 3D perception models that detect and assess infrastructure from LiDAR and imagery.
  • Develop LLM-powered products like agentic workflows, MCP servers, and AI copilots for customer tools.
  • Own models end-to-end, from data and training through evaluation, deployment, and monitoring in production.

Cyvl is a Physical AI company building purpose-built sensors, computer vision, and AI to turn every drive into current data for infrastructure. We're a fast-moving Boston startup with 500+ cities using our platform, and a culture of ownership, intensity, and care.

$150,000–$220,000/yr
Global Unlimited PTO

  • Lead the effort to make Runpod the fastest and most cost-efficient place for LLM inference, owning performance end to end.
  • Profile and diagnose performance bottlenecks across the serving stack, from scheduling to kernels, and implement fixes.
  • Work closely with product and infrastructure teams to shape how inference is offered, turning improvements into production-ready defaults.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We're a small, remote-first team that takes ownership seriously, moves fast, and has processed more than 20 billion inference requests.

  • Own large problems end to end across the stack, building a multi-agent orchestration SDK and CAD conversion pipeline.
  • Develop 3D-to-3D search across real customer part libraries and integrate constraint solvers with AI agents.
  • Address open problems in geometry reasoning and deploy hybrid cloud solutions for manufacturers.

3ive builds an AI co-engineer for physical products that works natively in 3D and learns from engineering decisions. The company is in production with major European manufacturers and has real revenue and customers.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

Global

  • Develop and evaluate deep learning models for feature detection, matching, depth estimation, and visual localization.
  • Own data preparation, training workflows, and evaluation of learned components for vision-based navigation.
  • Partner with state estimation and deployment engineers to integrate and optimize models for onboard systems.

Shield AI is a venture-backed defense-tech company developing intelligent systems to protect service members and civilians, with products like Hivemind autonomy software and V-BAT aircraft. It has offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, actively supporting operations worldwide.

Global

  • Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
  • Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
  • Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.

Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.

Europe

  • Own end-to-end technical strategy for major areas of the Voyager SDK’s application layer, including model deployment and pipeline development.
  • Set technical direction for image processing, create reusable architectures, and partner with the compiler team to resolve compilation issues.
  • Represent the Applications team’s capabilities to business stakeholders, mentor engineers, and drive documentation standards.

Axelera AI is a deep-tech company creating a next-generation AI platform to advance humanity. With $450 million raised and 250+ employees including 60+ PhDs, they have a strong business pipeline exceeding $100 million and are committed to innovation.

Europe

  • Lead in-depth technical discovery with engineering teams and customer stakeholders to understand AI inference requirements.
  • Translate customer objectives into production-ready architectures and define PoC success criteria.
  • Identify recurring workload patterns and communicate insights to Product and Engineering for platform evolution.

They are a partner company focused on AI infrastructure and performance-sensitive AI inference workloads. They have an international, engineering-led team solving complex challenges at the forefront of AI.

Global

  • Embed with customers to understand workflows, constraints, and success measures, then turn operational problems into technical plans.
  • Design, build, and deploy AI agents that integrate with customer tools, data, and business processes from prototype to production.
  • Measure real-world impact, iterate based on evidence, and share reusable patterns across deployments.

DehazeLabs transforms complex enterprise workflows into AI agents that operate in real-world environments. The company is a growing startup focused on applied AI, with a culture that blends engineering, product, and customer delivery.

US

  • Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
  • Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
  • Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality

Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.

$180,000–$250,000/yr
Global

  • Work directly with leading AI labs and enterprises to define research goals and technical requirements.
  • Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
  • Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.

Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.

$200,000–$350,000/yr
US Unlimited PTO

  • Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
  • Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
  • Lead technical story breakdowns, architectural design, and mentor engineers across the department.

Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.

Southeast Asia

  • Serve as the technical bridge between Tenstorrent and customers across Southeast Asia, guiding evaluations and deployments of AI platforms.
  • Troubleshoot hardware and software issues, and translate complex AI/ML concepts for both technical and non-technical audiences.
  • Build trust with customers and internal teams through proactive problem solving and regional travel.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. They value collaboration, curiosity, and a commitment to solving hard problems.

Spain

  • Design and develop large-scale platforms for LLM training and AI workloads.
  • Tackle distributed-systems challenges including intelligent job scheduling and resource optimization.
  • Collaborate with international teams to build production-ready AI infrastructure.

This role is with a partner company, an AI-focused R&D team building infrastructure for large language models. They are a fast-moving, highly technical team with a collaborative and innovative culture.

$86,400–$151,200/yr
Europe

  • Transform an existing AI agent product into scalable, reliable infrastructure.
  • Design and operate the core agent runtime, including orchestration, memory, and multi-agent workflows.
  • Work directly with the founder to set architecture, engineering standards, and development culture.

An early-stage AI infrastructure startup transforming a working AI agent product into scalable infrastructure. As the first engineering hire, you'll work directly with the founder to shape engineering culture and standards from the ground up.

$225,000–$250,000/yr
US

  • Design and implement state-of-the-art ML models and training pipelines for robotics.
  • Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
  • Collaborate with engineering to optimize training infrastructure and deployment.

We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.

Global

  • Design and implement novel training optimization techniques for large-scale neural networks.
  • Investigate approaches to improve training efficiency, stability, and convergence speed.
  • Collaborate with infrastructure and inference engineering teams to translate research into production performance.

This partner company focuses on AI research and training optimization for large-scale models. They maintain a small, senior-level team that values deep technical thinking and thoughtful execution.