Evaluate AI model responses on software engineering tasks using TypeScript.
Challenge AI systems across algorithms, data structures, and development practices.
Provide structured feedback to improve model reasoning and code quality.
You'll work on a cutting-edge AI training project where your TypeScript expertise directly improves advanced language models. The project is a flexible freelance opportunity with a focus on software engineering and coding challenges, though the size and culture of the hiring company are not specified.
Build and ship a project end to end: understand the problem, build a working version, and improve it based on user feedback.
Work with business teams to understand their workflows and translate requirements into actionable solutions.
Use AI-augmented development tools as your default workflow and develop judgment about their effectiveness.
M3 is a global healthcare technology company providing innovative research and technological solutions to the healthcare industry. The M3 Group operates in the US, Asia, and Europe with over 5.8 million physician members and is publicly traded on the Tokyo Stock Exchange, ranked in Forbes' Global 2000 list.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Participate in a paid remote video call about your AI-assisted development workflow.
Share how you use tools like GitHub Copilot, Cursor, or Claude Code to write and debug code.
Discuss strengths, limitations, and prompt strategies of current code generation models.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Ship real features in production using coding agents like Claude Code, Cursor, and Copilot, setting intent and owning the outcome.
Keep the codebase and context lean by maintaining standards, architecture notes, and feeding SAST findings back for real patches.
Build team-level tooling, review AI-assisted pull requests, and share workflows that improve how everyone works with agents.
Insider One is a marketing and customer engagement platform that unifies data, personalization, and journey orchestration across channels. The company is powered by 1500+ employees across 30+ offices, and is recognized as a woman-founded, women-led B2B SaaS unicorn.
Design, build, and ship custom internal AI tooling, agents, and workflows for autonomy and research.
Integrate and extend third-party AI tools and evaluate new AI models with structured pilots.
Partner with cross-functional teams to prototype solutions and drive adoption of AI tools.
Waabi is a leader in Physical AI, founded by AI visionary Raquel Urtasun. The company is growing quickly with offices in Toronto, San Francisco, Dallas, and Pittsburgh, and seeks diverse, innovative candidates.
Embed across teams to identify where AI can remove operational friction and build internal tools end-to-end.
Drive adoption through workshops, pairing, and office hours, measuring success by active usage.
Partner with Security to establish guardrails, access scoping, and audit trails for AI solutions.
LocalStack is a local cloud development platform that provides secure local sandboxes simulating cloud environments for developers and AI agents. They are a Series A startup with $25M funding, globally distributed teams across 25 countries, and a culture focused on ownership, openness, and excellence.
Debug and resolve issues across complex, multi-file codebases.
Implement new backend or full-stack functionality and refactor existing systems.
Develop and extend test suites, write technical specifications, and review code for quality.
Anyone AI is a company that recruits technical talent for project-based engagements with a leading AI lab. It offers flexible, remote consulting opportunities for experienced engineers.
Drive internal AI adoption across Customer Success, Sales, and Operations.
Build production AI agents and integrations using Python/TypeScript and LLM APIs.
Design human-in-the-loop workflows with guardrails and evaluation harnesses.
RoomPriceGenie is a hotel technology company founded in 2017 that provides an AI-driven pricing optimization solution. With a global customer base and a remote-first culture, the company is recognized as a Best Place to Work in HotelTech and values transparency, respect, and impact.
Assess software engineering tasks for technical accuracy, realism, and reproducibility.
Provide actionable feedback on codebase integration issues and logic errors.
Ensure AI training workflows are rigorous and practically applicable.
Project World Wide sources experienced technical specialists for AI training task auditing. This freelance contract opportunity focuses on ensuring technical rigor and accuracy in AI workflows.
Deliver end-to-end functionality across a TypeScript/React front end and a Golang backend.
Build products for partner integrations, including wallet systems and game systems in a certified environment.
Design and build APIs (GraphQL, gRPC, REST) and shape the architecture of a young platform on AWS.
Oddin builds products that partners integrate into their own platforms, focusing on AI-first development where 80-90% of code is generated with AI. They are a young startup with a collaborative, autonomous product team and no strict processes.
Embed with enterprise customers to prototype and ship production-grade AI-native applications and agentic workflows.
Advise engineering leadership on AI-native adoption strategy, tooling selection, and rollout sequencing.
Assess capability gaps and drive adoption of AI-native practices across the whole product engineering organization.
Astra provides an AI-native development platform and consulting services to help enterprises adopt AI-native software engineering. The company culture values AI-first, customer obsession, and continuous learning, with a focus on transforming engineering organizations.
Act as a hands-on technical partner for employees using Smart AI, helping them understand best practices and work through blockers.
Identify recurring needs and translate them into clear requirements for the Smart AI roadmap.
Host office hours, workshops, and create documentation to help employees build confidence with AI tools.
Coder is the leading platform for AI development infrastructure, enabling enterprises to securely run human and AI-driven development workflows. We are an equal opportunity employer committed to diversity and inclusion.
Code, test, and implement products for the HRS Platform Team with a focus on AI-driven engineering.
Integrate AI-assisted development tools like Claude Code and GitHub Copilot to increase velocity and quality.
Collaborate with teams, perform code reviews, and mentor others on effective and responsible AI tool use.
We build world-class products that simplify operations in Legal, Risk, Compliance and HR functions. We are a globally dispersed team with a diverse and inclusive culture.
Apply your expertise in software engineering to help train next-generation AI systems by creating reinforcement learning environments and solving complex software engineering problems.
Contribute expert-level code samples, debugging strategies, and development insights in languages such as Python, Java, Rust, Go, C++, or TypeScript.
Refactor and optimize code, review and validate peer contributions, and document technical reasoning to enhance AI training data quality.
The company is a rapidly growing, venture-backed AI company that combines world-class human expertise with advanced machine learning to build and improve cutting-edge AI models. It is backed by over $40 million in funding and has a rapidly expanding international network of experts.
Architect and ship AI-powered features that change how service businesses operate, creating net-new capabilities.
Build and maintain large-scale systems processing billions in payments and serving 120,000+ businesses with reliability.
Own your work end to end from technical design through production, with full accountability for performance.
Genius AI builds AI-powered products that handle admin work for local businesses, automating booking, marketing, client management, and finances. The company has saved businesses over 106 million admin hours and facilitated more than 624 million client touch points, fostering a culture of ownership and AI-first collaboration.
Design and implement full-stack product features across backend services and frontend interfaces, ensuring strong performance, reliability, security, and maintainability.
Take ownership of features from requirements through deployment and production support, using AI coding assistants to accelerate development.
Collaborate closely with product managers and engineers in an agile, feedback-driven environment focused on continuous improvement.
The company builds a modern data intelligence platform that enables organizations to discover, govern, and trust their data. They are a remote-first team with an agile, collaborative culture focused on continuous improvement and innovation.
Lead technical evaluations and proofs of concept for an enterprise AI coding platform in real customer environments.
Engage directly with CTOs, VPs of Engineering, and senior developers on AI agent architecture and integration.
Manage enterprise security and deployment reviews, troubleshoot integration issues, and feed customer insights into product.
The company is an enterprise AI coding platform provider helping engineering organizations accelerate development with AI agents and LLM infrastructure. It operates as a remote-first global team with a fast-paced, innovative culture focused on collaboration and rapid product iteration.
Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
Mentor and develop a team of engineers as the technical authority on agentic evaluation.
This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.
Build and maintain AI agent harnesses: tool calling, agent loops, context management, structured outputs.
Develop reusable tools and modules for AI across model types and sizes.
Design resilient systems for network failures, model errors, timeouts, and retries.
This company develops AI harness and peer-to-peer technology to enable local AI at scale. It is a fully remote, globally distributed engineering environment valuing autonomy, technical curiosity, and continuous innovation.