Utilize Automatic Prompt Generation (APG) tools to create baseline prompts for complex parent-child template clusters.
Monitor shadowbot runs and run prompt versions against gold data to measure autorater quality and accuracy metrics.
Manually draft, test, and refine prompts to navigate complex template architectures and handle edge cases.
Welo Data provides AI services specializing in data validation and LLM evaluation. They operate as a freelance-focused organization with a remote workforce, emphasizing technical agility and quality assurance.
Utilize automatic prompt generation tools to create baseline prompts for complex parent-child template clusters.
Run and supervise automated prompt optimization, review outputs, and flag deadlocks or plateaus.
Manually draft, test, and refine prompts to handle edge cases, anti-patterns, and solve complex template architectures.
Welo Data provides AI services focused on data validation and model evaluation. They operate remotely with a global team and emphasize technical expertise and collaboration.
Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
Develop prompt engineering strategies and evaluate model performance using quantitative methods.
Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.
GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.
Design and iterate complex system prompts and chain-of-thought structures for consumer AI experiences.
Partner with engineers and product managers to optimize prompt specifications for latency and cost.
Develop evaluation frameworks and playbooks to guardrail LLM outputs against bias and hallucination.
BOLD is a global company that creates digital products to help people build resumes, cover letters, and CVs, empowering job seekers in 180 countries. They are an established organization that values diversity and inclusion, with a culture of growth and professional fulfillment.
Design and implement production-ready AI capabilities that improve reasoning, accuracy, and explainability.
Optimize autonomous agent workflows and Retrieval-Augmented Generation pipelines for performance.
Develop evaluation frameworks and benchmark datasets to measure AI effectiveness and drive continuous improvement.
LTS builds an AI-native engineering platform for modernizing mission-critical healthcare systems serving millions of Veterans. The engineering team is intentionally small, giving every engineer meaningful ownership and direct influence over product direction.
Build and maintain MCP servers, tool definitions, and context management capabilities.
Design and implement intent evaluation, classification, and routing workflows.
Integrate Anthropic and OpenAI models into production environments.
Applaudo designs, builds, and scales AI-powered solutions that create real business impact. They foster a high-performance culture grounded in values such as empowerment, collaboration, and transparency.
Architect and build end-to-end GenAI applications using Python, LangChain, and LlamaIndex on Google Cloud.
Develop advanced RAG pipelines and Semantic Search systems for production-level accuracy.
Optimize LLM and Embedding fine-tuning while applying MLOps best practices for scalability.
Egen is a fast-growing and entrepreneurial company with a data-first mindset, using advanced technology platforms like Google Cloud and Salesforce to drive client impact through data and insights. We are dedicated to learning and innovation, with a culture that values engineering expertise and solving tough problems.
Lead end-to-end development for the AI platform and customer experiences, from initial roadmap to rollout.
Partner with engineers to design prompt strategies, evaluation frameworks, and guardrails balancing latency, cost, and accuracy.
Serve as the technical translator between engineering and the broader organization, establishing AI best practices and platform standards.
Jerry.ai is building the first AI agent to manage all your physical assets, starting with car insurance and expanding into home, motorcycle, and more. The company has 5M+ customers, raised $240M+, and has been profitable since 2024 with a fully remote team and offices in Palo Alto, New York, Chicago, and Toronto.
Define and lead prompt framework design and strategy, including architecture, version control, and deployment for distributed teams.
Design complex prompt systems like multi-agent and chain-of-thought architectures for resume builder flows across global product lines.
Partner with Engineering, Product, and Data teams to align prompt strategies with model selection, API cost reduction, and performance budgets.
BOLD helps people find jobs through digital products used in 180 countries. As an established global organization, they have empowered millions to build stronger resumes and cover letters, with a culture that values expertise, learning, and diversity.
Set technical direction for complex GenAI and agentic systems, building them hands-on and raising the bar for production-ready AI.
Own the hardest problems in AI portfolio, from agentic platforms at scale to AI in regulated, high-stakes environments.
Partner directly with client leadership to translate business strategy into AI architecture and shape solutions in pre-sales.
Egen is a fast-growing and entrepreneurial company with a data-first mindset, helping clients drive action and impact through data and insights using advanced technology platforms like Google Cloud and Salesforce. We are committed to being a place where the best people choose to work, dedicated to learning, thriving on solving tough problems, and continually innovating to achieve fast, effective results.
You will build AI-assisted tools, workflow automations, agents, prompts, and integrations to reduce manual effort and improve productivity.
You will partner with business stakeholders to understand high-friction workflows and deliver fit-for-purpose AI solutions.
You will implement engineering controls for data handling, access management, prompt safety, and output validation.
Shield AI is a venture-backed defense-tech company that develops intelligent systems, including Hivemind autonomy software and V-BAT and X-BAT aircraft, to protect service members and civilians. With offices across the U.S., Europe, the Middle East, and Asia-Pacific, the company's technology supports operations worldwide.
Design, develop, and deploy production-ready generative AI and machine learning applications.
Collaborate with cross-functional teams to integrate AI capabilities into business systems.
Apply prompt engineering, retrieval-augmented generation, and evaluation strategies to optimize AI performance.
The company is a technology firm specializing in AI transformation initiatives. It fosters a collaborative culture centered on learning, innovation, and professional development.
Evaluate AI quality across the advisor stack, including pre-call briefs, in-call guidance, and post-call outputs.
Iterate inside ORA by refining prompts, updating knowledge base entries, and tweaking skills to close the loop on issues.
Surface trends and drive continuous improvement by tagging conversations, logging issues, and recommending prioritized improvements.
HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.
Design and maintain LLM-powered backend services using Python and FastAPI.
Implement retrieval-augmented generation (RAG) for structured and unstructured fleet data.
Optimize retrieval accuracy, latency, and hallucination rates through automated evaluation pipelines.
Datakrew revolutionizes EV fleet intelligence with IoT and AI solutions. They aim to serve one million EVs within 5 years and cultivate a mission-driven culture.
Build and implement AI features by selecting the right model and approach for each use case.
Evaluate and monitor AI feature performance in production to ensure accuracy and reliability.
Improve and iterate on prompts and implementations based on how features behave in the wild.
Aphex is a construction execution platform that replaces traditional spreadsheets with collaborative tools for delivery teams. They are a remote-first company with a growing engineering team in the Philippines, serving major contractors on multi-billion dollar projects.
Build LLM-powered software by designing prompt flows and orchestrations to ensure great performance with no false positives.
Architect and build a production-grade, testable, and maintainable AI-powered software stack.
Design experiments and evaluation frameworks for system performance testing at scale, conducting data analysis to draw conclusions.
XBOW builds an AI-powered platform that autonomously discovers, validates, and exploits vulnerabilities, providing proof-backed results in hours instead of weeks. Founded by the creator of GitHub Copilot and backed by Sequoia and Altimeter, the team includes world-class AI engineers and security researchers.
Design, develop, and maintain LLM-powered applications and autonomous agents using Python and frameworks like LangGraph.
Build agentic workflows, implement prompt engineering strategies, and integrate LLM models across providers.
Develop observability tools, MCP Server components, and deploy cloud-native solutions on AWS.
NBCUniversal is one of the world's leading media and entertainment companies, creating and distributing content across film, television, and streaming, and operating theme parks. As a subsidiary of Comcast Corporation, the company champions an inclusive culture and strives to develop a talented workforce to serve communities.
Design and deploy state-of-the-art GenAI solutions including RAG pipelines, LLM orchestration, and AI agents.
Lead high-impact, client-facing engagements and oversee junior consultants.
Work hands-on with Databricks, LangChain, and models like GPT, LLaMA, and Claude.
Lovelytics helps some of the most complex enterprises modernize their data and put AI to work. The company has grown from 85 to 500+ people across the Americas, maintaining a technically excellent, low-ego culture.
Lead the delivery of AI projects end-to-end, establishing technical standards and mentoring a high-performing team.
Guide the design and delivery of RAG systems, agentic frameworks, and LLM-powered solutions for production.
Develop evaluation frameworks and quality standards to ensure reliable AI system performance.
Blend is a premier AI services provider that co-creates impactful solutions using data science, AI, and technology. The company fosters a culture of innovation and collaboration, with a focus on aligning human expertise with artificial intelligence.
Design, build, test, and iterate on prompts and prompt chains for production LLM workflows across client delivery and internal automation.
Implement evaluation harnesses, guardrails, and regression tests to ensure measurable and stable output quality.
Integrate LLM components with business systems via APIs, retrieval pipelines, and structured outputs.
Caravel is an award-winning NetSuite and Salesforce partner and part of BPM LLP, a Top-40 accounting and advisory firm. With 20+ years of experience and 1,000+ implementations, we are a fully remote team helping businesses across North America with honest technology solutions.