Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Apply quantization, batching, caching, and routing to optimize latency and cost at scale.
vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.
Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
Work across product, software development, performance engineering, and customer-facing implementations.
Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.
Develop and deploy machine learning and AI systems.
Work with LLMs, generative AI, and modern ML frameworks.
Optimize model performance, latency, and cost.
A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.
Lead the end-to-end lifecycle of language models and AI solutions, from research to production.
Navigate between closed and open-source ecosystems to maximize quality, optimize latency/cost, and ensure data governance.
Conduct applied research, experimentation, and data curation to create new AI capabilities and continuously improve existing ones.
Blip is a technology company that develops conversational AI and customer service platforms. The company fosters a culture of innovation and technical excellence, with a team of engineers and researchers working on cutting-edge AI solutions.
Own end-to-end delivery quality for major engagements, translating ambiguous client needs into practical execution plans.
Lead solution architecture and technical decision-making, making pragmatic tradeoffs between speed, quality, and client value.
Build and ship production AI/ML systems using Python, ML frameworks, and cloud-native infrastructure while mentoring other engineers.
Eliza is a technology services company and Advanced-tier OpenAI partner that helps organizations build and deploy AI solutions, from generative AI to predictive analytics. They are a collaborative, mission-driven team focused on real-world AI impact.
Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.
The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.
Architect scalable Python microservices and AI integrations on AWS.
Drive system health, code quality, and observability standards.
Mentor junior engineers and collaborate with US-based client teams.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. Our nearshore team culture empowers engineers to take ownership and grow continuously.
Develop and deploy production-ready AI solutions using Python, Azure, and LLMs.
Collaborate with cross-functional teams to translate business challenges into scalable applications.
Manage MLOps, CI/CD pipelines, and containerized deployments in an Agile environment.
Jobgether is an AI-powered job platform that matches candidates with roles using objective, fair reviews. It processes applications and shares shortlists with hiring companies, supporting a global remote workforce.
Taking ML or LLM proof-of-concept to production for large enterprises.
Designing and hardening data and training pipelines for enterprise ML systems.
Building LLM and RAG systems with retrieval quality, evaluation and cost control.
Janea Systems (USA) is a dynamic team of the best & brightest software engineering specialists and solutions innovators from around the world. From kernel to cloud, we provide high-impact software development services to Fortune 500 companies.
Design, build, and operate high-load distributed backend services powering the company's ML infrastructure.
Take end-to-end ownership of core ML services and data pipelines from design to deployment and continuous improvement.
Partner with ML and product teams to understand their needs and turn them into reliable, reusable platform capabilities.
Constructor is an AI-first e-commerce search and discovery platform that helps shoppers find products and enables brands to drive revenue. The company is fully remote, diverse, and values ownership and collaboration.
Design and develop AI-powered applications and enterprise solutions using Large Language Models.
Build scalable Python backend services and integrate AI with cloud platforms and databases.
Optimize AI model performance and collaborate with multidisciplinary teams.
The company is a partner organization that develops AI-powered solutions for enterprise clients. They have a diverse, international team and foster a culture of innovation and continuous learning.
Build AI-powered tools and copilots across the SDLC to reduce cognitive load and eliminate manual steps.
Research and deploy GenAI solutions to improve delivery pipelines and system reliability.
Collaborate with Platform, SRE, and DevOps teams to integrate intelligent automation into the core engineering platform.
Coderio designs and delivers scalable digital solutions for global companies. They combine strong technical expertise with a product mindset and value autonomy and clear communication.
Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.
Design and develop scalable backend services (Python, FastAPI) powering the AI platform.
Build and optimize RAG pipelines, AI agent systems, and orchestration of LLM models in production.
Manage cloud infrastructure (AWS/Azure) including multi-tenant isolation and serverless functions.
A dynamic startup building a modern B2B AI platform powered by autonomous agents. The company focuses on eliminating tedious analytics for enterprise teams through flexible architecture and advanced workflows.
Design and build backend services for AI-powered product features, including inference pipelines and orchestration layers around LLMs.
Develop high-throughput, low-latency distributed systems with monitoring, logging, and alerting across production services.
Collaborate with product, infrastructure, and AI engineers to optimize performance, caching, batching, and streaming.
This team is building an AI-native productivity platform that replaces repetitive digital work with reliable AI workflows. They are a small, focused product team working on cutting-edge AI infrastructure.
Spend your first weeks in the operator's seat, learning the customer's job from the inside before writing any code.
Ship production GenAI/LLM systems that move business unit metrics, not just complete scope.
Work embedded in small, senior teams alongside Principal Architects, owning the outcome from start to finish.
Provectus is a Premier AWS partner and an Anthropic Strategic Partner at the forefront of applied AI, helping enterprises turn Claude, agentic systems, and their own data into measurable business outcomes. With offices in North America, LATAM, and EMEA, we partner with clients worldwide and our team holds 100+ AWS certifications and is Claude Code certified.
Own Delivery End-to-End: Take ownership of features from design through to production.
Shape AI Agentic Systems: Build intelligent workflows using LLMs and agentic architectures.
We build production-ready AI agentic systems that reshape how financial institutions operate. We are a rapidly growing FinTech backed by major investors, recognized as Fintech of the Year for two consecutive years, with a team of experienced engineers and PhD-level AI specialists.
Build and own the model serving infrastructure, real-time inference, feature retrieval, and the latency budget that governs both.
Build the deployment path for data scientists to ship models, including bring-your-own-model support.
Own models in production: monitoring, drift detection, retraining, incident response, and the on-call rotation.
Sardine is the leading agentic risk platform for fighting financial crime. We are a remote-first company with hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo, hiring talented individuals with extreme ownership and high growth orientation.