Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.
Build the core AI platform that powers a copilot inside a new, AI-native trading experience.
Develop a high-performance backend in Rust supporting streaming responses, low-latency tool execution, caching, and reliable orchestration of model plus tools.
Design robust APIs and runtime components for AI capabilities including authentication, rate limits, auditing, tracing, retries, and fallbacks.
Clear Street is a diversified financial services firm replacing legacy infrastructure used across capital markets. Founded in 2018, the company has built a cloud-native clearing and custody system and supports billions in trading volume per day, with a culture focused on collaboration and diversity.
Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.
Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Own the architecture, design, and delivery for your team's platform area, turning roadmap into engineering plans that ship and hold up in production.
Design and build core parts of a compound AI system spanning agents, LLMs, and RAG, with a focus on reliability, evaluation, and scale.
Champion engineering best practices, mentor engineers, and partner with product and design counterparts to align on priorities.
Crogl builds fully autonomous AI team members for security operations centers (SOCs), helping organizations secure their digital estates. As a startup with a focus on innovation, they seek senior technical leaders to join their remote team.
Develop and operate production-ready AI and machine learning systems for enterprise-scale products.
Build and optimize LLM-powered applications, RAG pipelines, and intelligent agents.
Implement software engineering best practices for AI development including CI/CD and testing.
Our partner is building enterprise-grade AI solutions that deliver measurable business impact. They offer a remote-friendly work environment with a collaborative engineering culture focused on innovation, quality, and continuous learning.
Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
Architect scalable inference systems and collaborate closely with application engineering.
This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.
Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks.
Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement.
Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company offers innovative solutions with integrated intelligence on a single enterprise platform, connecting the pharmacy ecosystem.
Design, build, and deploy AI-powered solutions connecting Remote’s platform to customer systems and workflows.
Work at the frontier of practical AI, shipping reliable, observable, and secure systems in production.
Own customer outcomes from discovery through production rollout with high autonomy and direct influence on product roadmap.
Remote is a global HR platform that helps businesses recruit, pay, and manage international teams compliantly. The company fosters a future-focused, async culture with team members across 6 continents and emphasizes innovation, automation, and AI.
Evolving the AI knowledge platform with retrieval, indexing, and synthesis for organization-wide use.
Architecting and operating agentic infrastructure on AWS with cost guardrails and observability.
Partnering with product engineering to define the AI platform API surface and building reference agent implementations.
ShiftKey is a healthcare workforce marketplace that connects facilities with licensed professionals to fill shifts, addressing staffing shortages. The company fosters an inclusive and collaborative culture, valuing diverse perspectives.
Evolving the AI knowledge platform with retrieval, indexing, and synthesis layer for organization-wide use.
Architecting and operating agentic infrastructure on AWS for multi-step AI systems.
Designing graph-based retrieval across data sources for multi-hop queries.
ShiftKey is a healthcare workforce marketplace that connects facilities with licensed professionals to fill shifts, mitigating staffing shortages. The company is in a high-growth phase with a friendly and engaging work environment.
Own end-to-end technical execution for strategic customer and partner engagements, including discovery, infrastructure design, implementation, and production deployment.
Design and build cloud infrastructure supporting advanced AI workloads, including simulation, training, evaluation, inference, and large-scale batch processing.
Improve platform reliability, security, performance, and cost efficiency by debugging issues across application, network, storage, compute, and orchestration layers.
The partner company is building the infrastructure foundation for next-generation AI applications and physical AI workloads. The engineering team is pioneering and values ownership, technical excellence, and solving challenging engineering problems at scale.
Design and build a next-generation reliability platform for Affirm's production systems, blending distributed systems engineering with AI-assisted development.
Create AI agents and a centralized command center to assist with incident triage, root-cause analysis, and unified system health visualization.
Own projects end-to-end, from requirements to rollout, collaborating with partner teams to build powerful, simple solutions for developers.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later without hidden fees. The company is a remote-first organization with a strong focus on people-first values and inclusive benefits.
Build AI platform and production systems supporting computer vision, perception, simulation, and mapping products.
Own critical parts of the AI platform, including orchestration, APIs, and deployment patterns.
Partner with data scientists and simulation engineers to reduce iteration friction and design production architecture.
HERE Technologies is a location data and technology platform company that empowers customers to achieve better outcomes. The company is an equal opportunity employer with a focus on innovation and inclusion.
Lead the design and development of production AI systems powering insurance workflow automation.
Architect AI orchestration layers connecting LLMs, backend services, and business workflows.
Own end-to-end AI system design including inference pipelines, routing, caching, and fallback strategies.
BJAK is Southeast Asia's largest digital insurance platform, building AI-powered products that simplify insurance and financial services for millions of users. They use AI and automation to transform real-world insurance workflows and foster a high-ownership culture with small, empowered teams.
Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.
Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.
Build production-grade AI systems that enhance organizational intelligence and operational efficiency.
Collaborate with cross-functional teams to turn ambiguous AI opportunities into safe, measurable product capabilities.
Design and maintain scalable backend services, data pipelines, and human-in-the-loop workflows.
Honor Technology changes how society cares for older adults by providing technology, tools, and services for in-home care. With a global franchise network and over 100,000 Care Pros, Honor delivers 50 million hours of personalized care annually, fostering a culture of compassion and innovation.