Roles

Who Creates and Governs Your Synthetic Voices? Hire a Synthetic Voice Designer

A synthetic voice designer builds and maintains the AI voice clones a product or title speaks in, keeps the consent and licensing record behind every voice, and quality-controls output for naturalness and permitted use [1]. The role sits between audio engineering, casting and rights. Hire for the judgment about which take is publishable and which permission is missing, because the tool skill is the part that trains fastest.

The takeMost teams post this job as an audio job and screen it that way, then find out at launch that nobody can produce the release the cloned actor signed. The craft matters, but the craft is learnable in months and the paperwork instinct is not. The candidate worth hiring is the one who stops a session to ask who granted this voice and for what, before the render finishes. Screen for the stop, not the render.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing how they would handle a consent gap; it cannot capture them handling one. Olive puts that in front of them as work: an assignment, an assistant that will overreach, and a human reviewer who writes what actually happened at each moment.

Rank your shortlist

Start With the Take Where the Clone Slipped

The Spanish dub ships Thursday. The cloned lead sounds right for forty seconds, then reaches a line with a laugh inside it and produces something between a cough and a sigh. Nobody on the audio team can say whether to retune the model, book the human actor for a pickup, or ask first whether the licence even covers a laugh. That third question is the whole job.

That moment repeats across a surprising range of products. A banking app whose spoken alerts were recorded once in 2023 and now need eleven new phrases. A publisher pushing a backlist into audio. A game studio with a character whose actor is unavailable, or dead, or willing but expensive. In each case someone has to own both the sound and the permission, and in most companies nobody currently does.

Media job listings have started naming the role directly. Synthetic voice designer appears in the rewritten-jobs lists alongside a stated emphasis on ethical use and quality control, which is unusual language for a craft posting and tells you what employers got burned by 1. The alternate titles float around AI voice producer and voice clone manager; the title matters less than whether the person can hold both halves at once.

Write the job as one role even if you are tempted to split it. Split it and the audio side ships takes the rights side never cleared, and the rights side blocks takes that were fine.

What Separates a Synthetic Voice Designer From a Skilled Audio Editor?

The editor makes the take sound good. The designer decides whether the take should exist. In practice that means a person who can drive a voice model and also produce, from memory, the terms under which a given voice was licensed: which languages, which contexts, what happens on renewal, whether a synthesized read counts as a performance under the talent agreement.

The tells are specific. Ask a candidate to describe a voice they cloned and listen for what comes first. A performed answer opens with the model, the dataset size and how natural the result sounded. A real one opens with who the speaker was and what they agreed to, then gets to the audio. That ordering is hard to fake and easy to hear.

A second tell: ask what they cut. Everyone who has actually done this work has a story about a take that was technically fine and did not ship, usually because it put words in a person's mouth that the person had not agreed to say. Candidates without that story have run the tools but have not owned the decision.

A third: ask how they know a clone is degrading. The strong answer is procedural rather than aesthetic. They keep reference reads, they listen against the human original on a fixed set of hard lines, they track which phoneme contexts the model fails, and they can tell you the failure they see most often in their own work.

Watch for the opposite failure too. A candidate so cautious that no synthetic audio ever ships is not governing risk, only deferring it. You want the person who can say which uses are safe today, not only which ones frighten them.

Which Backgrounds Actually Produce This Person?

Post-production audio is the obvious pipeline, and it works: dialogue editors, ADR supervisors and dubbing mixers already spend their days matching performance across takes and languages. The less obvious sources are often better. Voice actors who moved into production understand consent from the side that grants it, and they negotiate with talent far more credibly than an engineer does.

Casting directors are a strong and underused source. The daily work is matching a voice to a role and reading a contract while doing it, which is most of this job with the model swapped for a person. Music licensing and clearance coordinators bring the second half almost fully formed and can learn the toolchain inside a quarter.

Speech researchers and ML engineers bring the deepest technical range and the largest gap. They can retrain and evaluate a model but often treat consent as somebody else's ticket. If you hire from here, pair the person with counsel early and check after ninety days whether the rights work is actually happening or quietly waiting.

On how the good ones got good with AI: nearly all of them taught themselves in public, on their own material. They cloned their own voice first because it was the only dataset they could consent to without asking. They kept a scratch corpus of hard lines, laughter, whispered reads, overlapping dialogue, numbers and names, and ran every new model against it. Ask for that corpus in the interview. The candidates who have one will send it before you finish the sentence, and the file itself tells you more than a portfolio reel does, because a reel shows the wins and the corpus shows what they were unable to fix. The same instinct shows up in adjacent AI-era roles, and it is what a good research agent orchestrator develops too: a private set of cases the tool keeps failing.

Where Do You Find One, and What Kills the Offer?

Look where dubbing and localization people already gather rather than in general AI channels. Localization and audio-post conferences, the SAG-AFTRA and Equity adjacent conversations about synthetic performance, audiobook and podcast production communities, and the credits of recent multi-language releases. Ask your existing dialogue editors and casting contacts who they call when a clone job comes in; this field is small enough that referral works better than search.

What closes this candidate is rarely money first. They want to know who has authority to overrule them. A designer who can be told to ship a voice they consider uncleared will leave, and they ask about that in the first conversation, sometimes obliquely. Give a direct answer: who signs off, what the escalation path is, and whether their objection stops a release or merely gets noted.

The second thing they care about is whether consent is real practice or decoration. Show them the actual talent agreement and the actual record of which voices are licensed for what. If that record does not exist yet, say so and make building it the first ninety days. Candidates respect the honest version and detect the polished one.

What kills offers, in rough order: no named legal partner, a manager who describes voice talent as a cost to be removed, a mandate to disclose nothing, and vague ownership of the model weights. On the identity side, verifying that a voice donor is who they claim to be borders on the work a candidate verification analyst does, and a candidate who has thought about impersonation risk will ask you how donors are authenticated before you ask them.

Budget the Role Before You Post It

No title-specific salary survey for synthetic voice designers existed at the time of writing, so treat any precise band you are handed with suspicion. The nearest public anchor is ZipRecruiter's posted AI voice job category, which as of 2026 advertises roughly $39 to $72 per hour across a mix of contract and staff work 2. That is a category range, not a band for this title, and it mixes voice talent with production roles.

Use it as a floor-finding device rather than a number. In practice teams hire this role in one of three shapes: a contract designer per title or per campaign, which fits studios with lumpy release schedules; a staff hire inside audio post or brand, which fits anyone with an ongoing voice product; and a split where an existing dialogue supervisor takes the craft and outside counsel takes the rights, which is cheapest and the most likely to fail quietly. Price the staff version against senior audio post roles in your market and add for the rights responsibility, because you are buying a second discipline.

On location: the craft travels and the equipment does not always. Editing, model work and QC are genuinely remote. Sessions with live talent, studio-grade recording of donor voices and anything under an embargo tend to pull people on-site, so most postings land at hybrid with named studio days rather than fully distributed. If you are remote-only, budget for travel to recording sessions instead of pretending they will not happen.

One more line item people forget: the consent record needs maintaining after launch, not only at signature. Someone has to re-check permissions at renewal, retire clones when a licence lapses, and answer a talent agent's question two years later. Whoever owns AI enablement across your teams, often an HR AI enablement partner in larger organizations, should know where that record lives before the first voice ships. Consent obligations for synthetic voice work vary by jurisdiction, union agreement and contract date, so check with counsel in the territories you publish in rather than relying on a general rule.

See how it works

Common questions

How do I become a synthetic voice designer?

Start from either half and add the other. From audio: learn a voice model toolchain on your own voice, build a corpus of hard lines including laughter, whispers, numbers and overlapping dialogue, and document every failure you cannot fix. From rights: if you already do clearance, casting or licensing, the paperwork half is done and the tools take a quarter to learn. Then get one real credit where you held both, even a small localization or audiobook job, and keep the consent record you built as your portfolio. Hiring managers ask about the take you refused to ship more often than the one you shipped.

What should a synthetic voice designer job description actually say?

Name both halves in the first line: builds and maintains AI voice clones for dubbing, ADR, character work or branded audio, and owns the consent and licensing record behind each voice 1. Then state authority explicitly, since candidates screen you on it. Say who signs off on a release, whether the designer can stop one, and who the named legal partner is. List the toolchain you actually use, the languages you ship in, and whether donor recording sessions happen on-site.

Can a dialogue editor or ADR supervisor do this job already?

Often yes for the craft, rarely yet for the rights. Post-production audio people already match performance across takes and languages, which is the harder technical instinct to teach. What they usually have not owned is the licensing side: what a given talent agreement permits, what happens at renewal, and when a synthesized read becomes a use nobody agreed to. If you promote internally, pair the person with counsel from day one and check at ninety days whether the consent record actually exists.

What are the legal consent requirements for AI voice work?

They depend on jurisdiction, union agreement and the date the underlying contract was signed, and they have been changing quickly. Older voice-over agreements frequently predate synthetic reproduction entirely, which means a licence to use a recording is not automatically a licence to clone the voice in it. Treat consent as specific rather than general: which voice, which uses, which languages, for how long, with what renewal terms. Get primary documents and check with counsel in every territory you publish in before shipping a cloned voice.

What does a synthetic voice designer cost to hire?

No title-specific salary survey existed at the time of writing. The nearest public anchor is ZipRecruiter's posted AI voice category, roughly $39 to $72 per hour as of 2026, which spans both talent and production work and is not a band for this title 2. Price a staff hire against senior audio post roles in your market and add for the rights responsibility. Contract-per-title is common at studios with irregular release schedules.

How do you test for judgment rather than tool skill in an interview?

Give a work sample with a defect planted in the paperwork rather than in the audio. Hand over a short dub brief, a talent agreement that does not clearly cover the requested use, and a model that produces an acceptable take. The candidate you want stops on the agreement before discussing the render, names precisely what permission is missing, and proposes a path that does not require guessing. The candidate to pass on delivers a clean file and a note about latency.

References

  1. 1. Media Industry Jobs Are Being Rewritten: This Is the New List Mediabistro, 2025. mediabistro.com Names synthetic voice designer as a rewritten media role covering AI voice clones for dubbing, ADR and character work, with stated emphasis on ethical use and quality.
  2. 2. AI Voice Jobs ZipRecruiter, 2026. ziprecruiter.com Posted US job category for AI voice work, advertising roughly $39 to $72 per hour as of 2026. A broad category spanning talent and production roles, not a band for this title.

2 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.