Speech Synthesis AI training jobs

Open AI training and model evaluation roles that call for Speech Synthesis expertise, gathered from 3 platforms. Part of Languages, Translation & Voice.

36
open jobs
$78/hr
median published rate
$280/hr
highest published rate
3
platforms hiring

Pay figures cover the 34 open jobs that publish an hourly rate in USD. They are published rates, not a guarantee of income. Last checked on October 8, 2026.

21 to 36 of 36 jobs
  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: Spanish

You will listen to audiobooks generated by text-to-speech synthesis in North American Spanish to evaluate the naturalness and quality of the narration. Your job is to detect pronunciation errors, inaccuracies, and acoustic defects, then document them precisely to improve the underlying models. This role suits a native Spanish speaker passionate about audiobooks and experienced with detailed annotation work.

Posted September 14, 2026

OpenRecently verified$15-20/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: xAI
  • Language: Bulgarian

This role involves training and refining an AI assistant's voice and multilingual audio abilities by annotating and curating speech data, with a focus on Bulgarian. It suits native Bulgarian speakers who have a strong ear for accents, pronunciation and audio quality and who can also localize English text into Bulgarian.

Posted August 7, 2026

OpenRecently verifiedPay not disclosed
  • Languages, Translation & Voice
  • Platform: xAI
  • Language: Khmer

This role asks a native Khmer speaker to annotate and record multilingual audio, and to translate English text into Khmer, so that an AI assistant handles speech and voice interactions well. It suits people with strong linguistic and auditory judgment who can work independently on ambiguous or noisy recordings.

Posted August 7, 2026

OpenRecently verifiedPay not disclosed
  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United Kingdom

You are a native English speaker based in the United Kingdom with a neutral and standard British English accent. You will record high-quality voice samples as part of developing advanced speech synthesis systems. Your recordings will be used exclusively for training a customer service AI agent and may be used for voice cloning.

Posted August 5, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Australia

You are a native English speaker with an authentic Australian accent and have professional or semi-professional recording equipment. This job involves recording high-quality voice samples on varied scripts for the training and evaluation of next-generation voice synthesis systems. Your voice will be cloned and used exclusively for a customer experience AI agent.

Posted July 30, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Switzerland
  • Language: German

This is a remote voice recording role for native Swiss German speakers who will read varied scripts to help build text-to-speech voices for an AI customer service agent. Experienced voice performers and newcomers with a clear, expressive voice and a suitable home recording setup are both welcome, and onboarding happens on a rolling basis.

Posted July 30, 2026

OpenRecently verified$50-150/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States

Are you a voice actor based in the United States with African American cultural background? This role invites you to record high-quality voice samples for training next-generation voice synthesis systems. Your recordings will be used exclusively to develop an AI conversational agent and may be used for voice cloning, under strictly confidential terms.

Posted July 17, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States

You are a native speaker of southern United States English whose clear and expressive voice must be recorded to train next-generation speech synthesis systems. This role involves participating in a recording session of approximately four hours, with the possibility of a second similar session, during which you will read various scripts while controlling intonation and emotions. Your recording will be used exclusively to develop a conversational AI agent for the client and may be used to clone your voice for this sole purpose.

Posted June 29, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Peru
  • Language: Spanish

This is a recurring voice recording engagement for building text-to-speech systems that sound natural and expressive. It suits native Spanish speakers with a clear regional accent, whether or not they have prior voice work. Each contributor records one session of about four hours, with a possible second session later.

Posted June 21, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: French

You record voice samples in standard French for the training of voice synthesis systems intended for an AI assistant. The role includes a recording session of approximately 4 hours, with the possibility of a second session depending on client needs. Your voice may be cloned for exclusive internal use.

Posted June 19, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States
  • Language: French

This role involves editing and validating French audio recordings for speech synthesis systems and conversational AI. You will prepare high-quality voice data corpora by performing audio cleaning, technical compliance checks, and linguistic evaluations. The ideal profile is an audio engineer with native or near-native French proficiency, with solid experience in voice editing and audio quality assurance.

Posted June 1, 2026

OpenRecently verified$50/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Canada
  • Language: French

This role consists of recording high-quality voice samples in Canadian French and neutral French for the training and evaluation of next-generation voice synthesis systems. You must be a native Canadian French speaker based in Canada, have a clear and expressive voice, and agree to voice cloning of your recordings for an internal AI agent. An approximately 4-hour recording session is planned, with the possibility of a second session if needed.

Posted June 1, 2026

OpenRecently verified$50/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Germany
  • Language: German

This job involves recording high-quality voice samples in German for training and evaluating next-generation speech synthesis systems. Your voice will be used exclusively for an AI customer service agent and may be cloned for this purpose. The role is aimed at native German speakers based in Germany, with a natural and expressive voice, and no prior professional experience is required.

Posted June 1, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: 18 countries
  • Language: Spanish

This opening seeks Spanish-speaking voice talent to record scripts that will help train a text-to-speech system for a customer service AI agent. The work is one recording session of about four hours, with a second session possible but not guaranteed, and the recordings are restricted to that single internal use. It suits both seasoned voice professionals and newcomers who have a clear voice and a solid recording setup.

Posted June 1, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: Spain
  • Language: Spanish

This job involves recording Castilian Spanish voice samples used to train and test text-to-speech systems that aim for natural, expressive speech. It suits native Peninsular Spanish speakers, whether experienced voice professionals or newcomers with a clear voice and a good recording setup. The work is a single session of about four hours, with a possible second session later.

Posted June 1, 2026

OpenRecently verified$50-100/hr

Referral link: we may earn a fee. Apply without it

  • Languages, Translation & Voice
  • Platform: Mercor
  • Location: United States

This role involves providing high-quality voice recordings for training next-generation voice synthesis systems. You will perform various scripts (conversational, narrative, educational) in a recording session of approximately 4 hours, with the possibility of a second session later. We are looking for a native American English speaker with a clear and expressive voice who accepts voice cloning for the needs of an internal AI agent.

Posted March 20, 2026

OpenRecently verified$50-150/hr

Referral link: we may earn a fee. Apply without it

See which of these jobs match your experience

Add your resume once. Every open job is scored against your profile, and we email you only when a strong match appears. Free.

Get recommended jobs

FAQ

Questions about Speech Synthesis AI training jobs

How many Speech Synthesis AI training jobs are open right now?

36 Speech Synthesis AI training jobs are open on SideHustler today. Listings are rechecked against the official postings: the latest check on this list was on October 8, 2026.

How much do Speech Synthesis AI training jobs pay?

Among the 34 open roles that publish an hourly rate in USD, pay runs from $15 to $280 per hour, with a median of $78. These are the rates the platforms publish, not a guarantee of income or hours.

Which platforms offer Speech Synthesis AI training jobs?

Alignerr (17), Mercor (17) and xAI (2). You apply on the platform itself, which handles screening, contracts and payment.

What skills do Speech Synthesis AI training jobs ask for most?

The most requested areas of expertise on these listings are Speech Synthesis, Voice Recording, Voice Acting, Multilingual Expertise, Language Evaluation. Each listing states its own requirements.

How do I apply?

Apply now opens the official posting on the platform, where you complete the application. You need no account on SideHustler. To see which roles fit your background first, add your resume and every open job is scored against your profile.