Conduct groundbreaking research in generative audio using diffusion or flow matching models, focusing on vocal synthesis, post-training, or editing.
Run large-scale experiments with access to Spotify’s extensive infrastructure and audience.
Collaborate with cross-functional teams to craft innovative solutions and publish findings at top conferences.
Spotify is a global music streaming platform with over 700 million monthly active users. We are an equal opportunity employer committed to inclusivity and innovation.
Improve core inference services including networking, speech processing, audio transcoding, and latency optimization.
Develop processes for measuring, building, and optimizing services to maximize system performance.
Debug complex system issues involving networking, scheduling, and high performance computing interactions.
Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text, text-to-speech, and voice agents. With over 200,000 developers and 1,300+ organizations, including major partners like Twilio and Cloudflare, Deepgram emphasizes an AI-first mindset and rapid innovation.
Review pre-labeled audio segmentation and transcription for accuracy.
Follow project guidelines to ensure high-quality results.
Contribute to improving AI speech recognition models.
Appen is a global leader in providing AI training data and services. They have a large, diverse team of independent contractors and a culture focused on flexibility and innovation.
Transcribe audio recordings accurately into text, capturing speech including pauses, filler words, and speaker labels.
Add timestamps and mark non-speech events such as laughter or background noise.
Review and refine transcripts to ensure high-quality data for AI speech recognition models.
Appen provides high-quality training data for AI and machine learning systems. It operates a global platform with a large community of independent contractors, offering flexible remote work opportunities.
Record natural, conversational customer-service style audio from your home recording setup.
Submit audition recordings and complete two supervised sessions of up to 4 hours each.
Use a professional microphone and quiet space to meet technical audio specifications.
Welo Data provides AI services, including data collection and annotation for conversational AI applications. They operate as a freelance platform with a global community of contributors, focusing on high-quality data generation.
Listen to audio recordings in your native language and transcribe them verbatim.
Follow provided transcription guidelines to deliver high-quality, accurate transcriptions.
Help create ground-truth data used to evaluate and improve automatic speech recognition systems.
RWS is a technology and language services company specializing in AI training data and automatic speech recognition. They are a large global organization that embraces DEI and promotes equal opportunity, offering flexible freelance work.
Drive product strategy and roadmap for audio advertising tools, blending user needs with business goals.
Collaborate with engineering, design, and sales teams to deliver scalable, impactful features.
Analyze market trends and customer feedback to prioritize features and optimize product performance.
Audiohook is a fast-growing technology company transforming the way brands connect with audiences through audio. As a dynamic, fully remote team, we build innovative tools that help advertisers reach listeners across streaming platforms, podcasts, and emerging audio channels.
Record short video and/or audio clips in German with various emotional tones.
Listen to and transcribe video and/or audio clips accurately.
Work remotely on a flexible schedule with a quiet, distraction-free recording environment.
The enterprise client is involved in AI research aimed at helping AI better understand human tone, intent, and emotion. The team is looking for expressive individuals to contribute to video and audio tasks.
Compose original scores, themes, songs, and background music for animated and live-action productions.
Edit, arrange, and place music within completed episodes to enhance pacing, emotion, and narrative impact.
Experiment with AI music-generation workflows to rapidly prototype and refine compositions.
TrueShort is one of the fastest-growing streaming apps in the United States, operating as an AI studio and streaming platform that experiments, develops, and releases original series and films constantly without development purgatory. Backed by prominent investors including Jeffrey Katzenberg and Khosla Ventures, the team is a small, fast-moving crew working at the frontier of AI-powered entertainment production across animation, live action, and emerging formats.
Record 150 short, scripted voice prompts in Italian using the Appen Mobile app on your smartphone.
Work from anywhere in Italy on your own schedule, with the task taking approximately 2 hours.
Must be a native Italian speaker, currently residing in Italy, and 18+ years old with a compatible smartphone.
CrowdGen by Appen is a platform that connects independent contractors with AI data collection projects. They are a large company focused on gathering high-quality speech data to improve voice AI systems, with a culture of flexible, project-based work.
Perform voice recordings for assigned projects with clarity, accuracy, and professionalism.
Take direction and incorporate feedback from team members and project leads.
Participate in recording sessions in studio or remote environments using approved recording software and equipment.
Ibility is a Service-Disabled Veteran-Owned and Woman-Owned Small Business that helps government leaders achieve their mission through human-centered design products and programs. The small but mighty team is fun, passionate, bold, and creative, positioned for rapid growth.
Work directly with customer engineering teams to design and implement real-time applications on the LiveKit platform.
Architect scalable solutions for voice AI and developer platforms, building prototypes and reference implementations.
Debug production issues, lead technical workshops, and translate customer needs into product improvements.
LiveKit builds infrastructure for the agentic era of computing, enabling developers to build, test, deploy, and scale AI agents in production. Founded in 2021, the company powers voice and agentic AI for leading enterprises like OpenAI, Salesforce, and Meta, with a focus on innovation and collaboration.
Curate and annotate multilingual audio data to train AI for voice interactions and speech recognition.
Ensure high-quality voice recordings and accurate transcriptions across diverse languages and accents.
Collaborate with technical staff to improve annotation tools and audio workflows.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence.
Respond quickly to inbound leads from marketing campaigns and product sign-ups, qualifying them into sales opportunities.
Manage and prioritize high volume of inbound requests to ensure timely responses.
Work with Account Executives and fellow SDRs to hand off qualified opportunities and exceed quarterly targets.
Deepgram is a leading Voice AI platform providing APIs for speech-to-text, text-to-speech, and voice agents at scale. They serve over 1,300 organizations and 200,000 developers, backed by Series C funding, with a fast-paced, AI-first culture.
Architect and re-engineer audit workflows by applying Lean/Six Sigma to eliminate waste and identify automation opportunities.
Orchestrate AI agents and build prompt libraries to automate control mapping, walkthroughs, and anomaly narrative generation.
Design agentic automations and real-time risk dashboards using Power Automate, DataSnipper, and Power BI.
IonQ is the world's leading quantum platform and merchant supplier, delivering integrated quantum solutions for computing, networking, sensing, and security. With operations across multiple countries and a world record in quantum computing performance, the company fosters a culture of autonomy, productivity, and respect.
Record natural, conversational customer-service style speech for a conversational AI application.
Submit an audition consisting of one unscripted and one scripted sample.
Complete two supervised recording sessions of approximately 4 hours each using your own professional home studio.
Welo Data provides AI services including data collection and annotation for conversational AI applications. They collaborate with a global network of freelancers and prioritize high-quality data for AI training.
Own the voice pipeline end-to-end, including ASR, TTS, streaming, and vendor relationships.
Set strategy and roadmap for voice product surface, balancing customer commitments and technology bets.
Work directly with engineering leads, customers, and executives to drive outcomes.
ASAPP delivers the best AI-powered customer experience. They are a globally diverse team with hubs in New York City, Mountain View, Latin America, and India, operating in a fast-paced, high-growth startup environment.