Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.
Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.
Lead in-depth technical discovery with engineering teams and customer stakeholders to understand AI inference requirements.
Translate customer objectives into production-ready architectures and define PoC success criteria.
Identify recurring workload patterns and communicate insights to Product and Engineering for platform evolution.
They are a partner company focused on AI infrastructure and performance-sensitive AI inference workloads. They have an international, engineering-led team solving complex challenges at the forefront of AI.
Operate and expand Telnyx's own B300 GPU fleet to maximize inference throughput per GPU-dollar.
Design and implement serverless inference for open-weight models and dedicated enterprise deployments.
Work upstream in open-source technologies like vLLM, SGLang, and Kubernetes.
Telnyx is an industry leader building the future of global connectivity through a private, multi-cloud IP network and edge APIs. The company is financially stable and profitable, with a global team and a focus on innovation and continuous learning.
Architect and own the technical roadmap for a secure local LLM platform deployed in company-controlled infrastructure.
Build modular inference layers with stable APIs, model routing, and production-grade serving optimizations.
Design and operate retrieval-augmented generation pipelines with permission-aware access and systematic evaluation.
Parallel Wireless is a U.S.-based pioneer in Open RAN innovation, transforming how mobile networks are built and powered. The company is a leader in software-centric, hardware-agnostic network solutions with a focus on reducing complexity and total cost of ownership.
Design, build, and ship production services, APIs, and user-facing interfaces.
Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.
Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.
Design, build, and deploy production ML and LLM-based systems for enterprise clients.
Own technical delivery end-to-end from architecture to deployment and iteration.
Mentor and support other ML engineers through code reviews and technical guidance.
TensorOps is a boutique AI consultancy that designs and ships production-grade AI systems for enterprise clients. The company has shipped AI systems impacting 200M+ end users daily, partnered with 11 unicorns, and operates fully remotely with a supportive, fast-growing culture.
Evaluate GPU kernel tasks for technical accuracy, realism, solvability, reproducibility, and robust testing criteria.
Rigorously test and troubleshoot complex GPU programming scenarios to identify memory allocation bugs, execution bottlenecks, and parallel computing logic errors.
Review CUDA, Triton, and other GPU kernel implementations and provide clear, actionable technical feedback.
Our partner is a company specializing in AI training and evaluation, seeking experienced GPU kernel specialists to audit AI training tasks. The project is globally distributed and offers fully remote freelance work.
Design and build ML infrastructure and GenAI models for voice and speech, including LLMs, RAG, and real-time audio at scale.
Partner with engineering teams to integrate AI capabilities into customer-facing products, ensuring scalable and reliable solutions.
Mentor engineers, explore emerging AI tech, and contribute to a collaborative, autonomous, and innovative culture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company focuses on using technology to streamline recruitment processes and ensure fair candidate evaluation.
Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
Lead technical story breakdowns, architectural design, and mentor engineers across the department.
Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Own the ML strategy for dialogue systems, lead a team of 3 ML engineers, and drive LLM post-training and model adaptation.
Build evaluation layers, cut dialogue failure modes, and keep inference efficient on latency and cost.
Stay hands-on with prototypes, debugging agent traces, and reviewing team work.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness and isolation through social entertainment platforms. The company has an international team of digital nomads and has been recognized as a Great Place to Work winner and a top company for remote jobs.
Build and maintain core infrastructure for Quora's ML platform, ensuring high availability, scalability, and performance.
Build and improve distributed systems serving ML models in production, from Large Recommendation Models to Large Language Models.
Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.
Quora's mission is to grow the world's collective intelligence through two platforms: Quora for global knowledge sharing and Poe for AI agent collaboration. We are a remote-first company with passionate, collaborative, and high-performing global teams, rooted in transparency and experimentation.