Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.
Incident ManagementSite Reliability EngineeringReliability EngineeringCritical EvaluationWritten Communication
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Assess software engineering tasks for technical accuracy, realism, and reproducibility.
Provide actionable feedback on codebase integration issues and logic errors.
Ensure AI training workflows are rigorous and practically applicable.
Project World Wide sources experienced technical specialists for AI training task auditing. This freelance contract opportunity focuses on ensuring technical rigor and accuracy in AI workflows.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Evaluate AI-generated documents, spreadsheets, and presentations against domain-specific quality standards.
Assess outputs for factual accuracy, procurement relevance, completeness, clarity, and practical applicability.
Provide structured written feedback identifying issues and opportunities for improvement.
This partner company specializes in evaluating AI-generated work products for public-sector procurement and RFI responses. The company size and culture are not specified in the posting, but the engagement is flexible and fully remote.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Manage team performance, career development, and project prioritization while driving a culture of automation.
Drive initiatives with partner teams to improve infrastructure reliability and act as crisis management.
Analyze existing processes to drive continuous improvement and efficiencies.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter with an intelligent cloud platform. We are building an AI-native culture where technology and talent are unstoppable together, serving over 8,100 customers.
Audit AWS Serverless and Infrastructure as Code tasks for technical accuracy and realism.
Evaluate deployment scenarios, architecture logic, and testing criteria.
Provide clear, actionable feedback to improve AI training and evaluation systems.
Jobgether uses AI-powered matching to connect professionals with freelance roles at partner companies. This project offers independent remote work with competitive hourly rates and flexible scheduling.
Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
Contribute to improving AI systems by providing expert judgment and structured feedback.
The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.
Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Evaluate AI-generated market research and competitive intelligence artifacts for accuracy, rigor, and quality.
Apply structured rubrics to assess deliverables and identify factual inaccuracies and analytical gaps.
Provide clear, actionable written feedback to support evaluation decisions and improve AI training.
The partner company specializes in evaluating AI-generated market research and competitive intelligence content. They hire independent contractors for flexible remote engagements.
Evaluate AI-generated work products in real estate, hospitality, and events using quality rubrics.
Identify factual, aesthetic, and presentation errors and provide actionable feedback.
Apply industry expertise to distinguish realistic, commercially sound work from generic AI content.
The company develops AI systems and evaluates their outputs for quality. They seek experienced industry professionals for flexible remote contract work.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Assess technical accuracy and reproducibility of AWS Trainium/NKI tasks.
Provide actionable feedback on kernel execution bugs and logic errors.
Evaluate hardware acceleration inefficiencies and compilation issues.
We source experienced technical specialists to audit AI training tasks and evaluation workflows. The project is remote and freelance, with a focus on technical accuracy and efficiency.
Collaborate with engineers and clients to capture technical requirements, decisions, and tradeoffs.
Track project plans, risks, dependencies, and deliverables across AI infrastructure, data, and security.
Build trusted relationships with client stakeholders and ensure alignment with business objectives.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots in the open-source ecosystem, with a culture of ownership and collaboration.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.