Similar Jobs
See allLLM Inference Engineer
NEAR AI
US
PyTorch
CUDA
Staff Software Engineer, AI Inference
Syllo
US
Python
Go
Rust
GPU Performance and Benchmarking Engineer
Vultr
Global
Python
Senior/Staff Engineer – Computing Architecture
Axelera AI
Europe
Python
C++
Performance Analysis
Lead Machine Learning Engineer, Inference & Performance
Egen
Python
Kubernetes
VLLM
Responsibilities:
- Develop and optimize low-level kernels, runtime components, and system software responsible for high-performance AI inference workloads.
- Improve inference engine performance across GPU platforms by identifying bottlenecks and implementing advanced optimization techniques.
- Profile, debug, and resolve system-level and hardware-level performance issues across CPU and GPU environments.
Requirements:
- Strong proficiency in C++ development or deep expertise in GPU programming focused on low-level, high-performance computing and memory management.
- Experience with GPU programming or systems-level software development, including operating system internals, kernel modules, device drivers.
- Hands-on experience using profiling and debugging tools to analyze CPU and GPU performance issues and optimize code based on findings.
Benefits:
- Competitive compensation package.
- Fully remote working flexibility within Europe.
- Opportunity to work on impactful AI infrastructure projects at global scale.
Jobgether
This position is listed on behalf of a partner company building cutting-edge AI infrastructure for large-scale inference platforms. They operate in a highly technical, international, and innovation-driven environment where engineering excellence and ownership are valued.