Source Job

North America

  • Analyze and optimize CPU cluster performance including cache hierarchies and interconnects.
  • Build and use performance models and simulation environments to evaluate architectural concepts.
  • Lead architectural tradeoff studies across performance, scalability, and power to influence design decisions.

Simulation

5 jobs similar to Performance Architect, CPU Cluster

Jobs ranked by similarity.

North America Spain

  • Verify the fetch unit for next-gen high-performance RISC-V CPU cores, including instruction fetch, prediction, decode, and dispatch.
  • Develop verification strategies, create sophisticated tests, and debug complex failures using SystemVerilog and UVM.
  • Collaborate with CPU architects and RTL designers to ensure front-end correctness across workloads and corner cases.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible.

North America Spain

  • Work closely with CPU architects, RTL designers, and software engineers to develop verification strategies.
  • Build robust test environments and uncover corner cases in modern out-of-order processors.
  • Own meaningful portions of the verification effort from test plans to debugging failures.

Tenstorrent is a leader in AI technology, developing high-performance RISC-V CPUs and AI platforms. The company values collaboration, curiosity, and solving hard problems, and is growing its team with contributors of all seniorities.

Global

  • Evaluate GPU kernel tasks for technical accuracy, realism, solvability, reproducibility, and robust testing criteria.
  • Rigorously test and troubleshoot complex GPU programming scenarios to identify memory allocation bugs, execution bottlenecks, and parallel computing logic errors.
  • Review CUDA, Triton, and other GPU kernel implementations and provide clear, actionable technical feedback.

Our partner is a company specializing in AI training and evaluation, seeking experienced GPU kernel specialists to audit AI training tasks. The project is globally distributed and offers fully remote freelance work.

Canada

  • Lead the architecture and delivery of a large-scale GPU infrastructure platform, evolving from managed Kubernetes to bare-metal with Slurm and inference support.
  • Manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting technical standards and overseeing implementation.
  • Own GPU infrastructure operations, including Slurm, Kubernetes, NVIDIA hardware, observability, and incident response, while acting as the primary technical interface with partners.

Jobgether is an AI-powered job matching platform that connects candidates with relevant roles, ensuring a fair and objective review process. It operates with a distributed team and partners with companies globally, focusing on efficient and transparent recruitment.

US

  • Own the technical architecture and roadmap for a multi-RAT RAN digital twin covering LTE and 5G NR.
  • Integrate production MAC and scheduler software into deterministic closed-loop simulations with stable interfaces.
  • Develop AI/ML-based RAN capabilities including neural channel estimation and learned scheduling policies.

Parallel Wireless is a U.S.-based pioneer in Open RAN innovation, transforming how mobile networks are built, optimized, and powered. The company culture emphasizes software-centric, hardware-agnostic approaches to reduce complexity and total cost of ownership.