Roles
Who Owns Content Understanding When a Catalog Is Too Large to Watch?
A content understanding AI product manager owns what machines say about a catalog: the tag taxonomy, the model output that fills it, the confidence threshold at which a tag publishes without a human, and the correction path when a tag is wrong on a title that millions can see. The role sits between metadata operations and applied research, and Netflix is hiring it now under the title AI Product Manager, Catalog and Content Understanding [1].
The takeMetadata used to be a back-office ledger, so it was staffed like one. The moment a model started generating the ledger, the interesting decisions moved: what the taxonomy is allowed to assert about a person on screen, what publishes unreviewed, who is accountable when a title is mislabeled in a way an audience notices. Those are product decisions with editorial and legal weight, and they are still landing on operations leads who were never given the authority to make them. Give the decisions an owner before an embarrassing tag makes the case for you.
Where Olive fits
Open a role and see what the work shows
No screen can tell you which resume a model wrote, so Olive skips the artifact and assesses the person: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six findings with the timestamp behind each one. The candidate gets the same report you do.
Rank your shortlistA Model Tagged a Documentary as a Comedy and Nobody Could Say Whose Call That Was
The tag surfaced on a home row, a viewer screenshotted it, and by the afternoon three teams were in a thread. Research owned the model. Operations owned the taxonomy. Recommendations consumed the field and had already trained on a quarter of it. Nobody owned the sentence that would have prevented it, which is the sentence about what confidence level lets a genre tag publish without a person reading it. That gap is the job.
A content understanding AI product manager owns what the machine asserts about a piece of content. The deliverables are concrete: a taxonomy with definitions precise enough to be scored, an evaluation set of titles with known-correct labels, a publish policy tied to confidence and to blast radius, and a contract with every downstream consumer about what a field means and how wrong it can be.
The first tell of a real one is that they treat the taxonomy as a specification rather than a list. Ask a candidate to define a label like dark or slow-paced well enough that two annotators would agree. Weak answers give a synonym. Strong answers give a definition, then immediately name the boundary case that breaks it, then say how they would resolve it and write the resolution back into the guideline. That loop is the craft.
The second tell is that they think in blast radius rather than in accuracy. A mistaken supporting-cast credit and a mistaken content warning are both single wrong fields and are not the same event. Listen for a candidate who sorts fields by consequence before they talk about model quality, and who can say which labels should never publish unreviewed no matter how confident the model is.
The third tell is that they know evaluation on a catalog is a sampling problem. You cannot audit millions of titles, so quality is a stratified sample with a real design behind it, weighted toward the titles that get watched and the labels that hurt. A candidate who proposes to review the whole thing, or who proposes to trust an aggregate accuracy number, has not run this at catalog scale.
The performed version of this role talks fluently about multimodal embeddings and shot detection. The real version has an opinion about who gets to add a label to the taxonomy, and can tell you about a time they refused one.
Which Backgrounds Actually Produce Someone Who Can Specify a Taxonomy?
The dependable feeders are people who have already been accountable for a field that other systems consume: catalog and metadata leads at a studio or distributor, search and ranking product managers, and data product managers who shipped a schema with real customers on the other side. All three have lived with the fact that a definition change is a migration, not an edit, and that somebody downstream is already depending on the old meaning.
Metadata operations leads convert best and get overlooked most. They wrote the annotation guidelines, arbitrated the disagreements, and know exactly which labels have never been applied consistently by anyone. What they usually lack is model literacy, which is a few months of work; what they carry is a decade of knowledge about where the taxonomy is quietly broken, which cannot be hired any other way. The same argument for promoting the person who already owns the data shows up in the digital twin data quality specialist hire.
The unexpected sources are worth a deliberate look. Accessibility and localization leads have spent careers writing descriptions of video that must be accurate and are read aloud to someone, which is content understanding with a stricter standard than any recommendation feature applies. Standards and practices reviewers know what a content warning is legally and editorially doing. Archivists and librarians have formal training in controlled vocabularies, which is the discipline this role keeps reinventing badly. Trust and safety classification leads bring the instinct for threshold policy, close enough that the safeguards product manager pipeline overlaps.
Two profiles interview well and struggle. Research-adjacent candidates optimize the model and treat the taxonomy as fixed input, which inverts the job. And pure consumer feature PMs can specify the surface beautifully while having no view on what the field underneath is allowed to claim. Ask both of them what they would delete from an existing taxonomy. The answer separates them quickly.
Rights and licensing fluency is a genuine bonus, since what a platform may assert or display about a title is often contractual rather than technical. Candidates from the AI data partnerships manager world arrive with that already.
Ask How the Candidate Got Good at Not Believing a Confident Label
Ask how they personally learned to work with these models and listen for a specific failure they own. The useful answer names a moment: a caption that read perfectly and described a scene that was not in the title, a summary that invented a character name, a tag pipeline that looked excellent in aggregate and was systematically wrong on one genre. Then they name the check they now run every time.
The answers worth hearing tend to be unglamorous. Someone watched twenty titles themselves against the model output before approving a pipeline, because reading the metrics was not the same as reading the catalog. Someone else discovered a quality number was carried entirely by easy titles and rebuilt the sample so hard cases were not diluted away. A third keeps a running file of cases where a model produced something plausible and wrong, and that file became the first evaluation set the team ever had.
The habit underneath all three is checking a confident claim against the source. A model says a scene contains a specific object at 00:14:22, and this person opens the title at 00:14:22. That is not a technical skill and it does not appear on a resume. It shows up in fifteen minutes of real work and almost nowhere else.
So make the interview a working session rather than a conversation. Hand over fifty model-generated tags for ten titles, with six errors you have already found, and give ninety minutes with whatever assistant they want. Ask for one page: which errors they found, which they would have caught with a policy rather than a review, what they would change in the taxonomy, and which labels they would stop publishing automatically. You will learn more from that page than from three rounds of discussion about multimodal architecture.
Be explicit that using an AI assistant during the exercise is expected. This role is done with these tools all day, and watching someone push back on an assistant that overreaches is the closest you will get to watching them do the job.
Where Do You Find This Person While the Title Is Still Forming?
Start with the employers visibly hiring the scope. Netflix currently lists an AI Product Manager for Catalog and Content Understanding in Los Gatos, pairing the metadata surface with model-driven understanding of video, work once split between operations staff and a research team 1. Streaming platforms, music and podcast catalogs, stock media libraries and large marketplaces all have this problem, and most have not named the role. Search by responsibility, not by title.
Look inside your own building first, because the category is young enough that external supply is thin. The person who owns your annotation guidelines, the localization lead, the search relevance analyst who keeps filing tickets about bad genre data: one of them is already doing the unpaid version of this job. Promoting them and buying the model literacy is usually faster than a six-month external search, and the internal candidate already knows which parts of the taxonomy are fiction.
What closes these candidates is authority over the taxonomy and over the publish threshold. They will ask who can overrule them when a partner wants a label added, whether they own the evaluation set or borrow it from research, and whether they can stop a rollout. Answer all three plainly. If the honest answer is that research owns the threshold, say so in the interview rather than in month two.
The second lever is showing them the mess. Bring a real example of a bad tag on a real title into the process. Candidates worth hiring read that as a team willing to look at its own errors, which is the working condition they are actually shopping for. Teams that present a tidy version of the problem lose these people to teams that do not.
What Should This Role Cost, and Does It Need to Sit Near the Catalog?
No wage series covers this title and no survey found for this piece prices it, so treat any point estimate as a guess wearing a benchmark's clothes. Hire against your senior product manager band, or principal if the role carries authority to hold a release, and check that band against local senior PM comp where the job sits. Catalog scale should move the number, not the word AI in the title.
The pressure on that band is real even where a specific figure is not. PwC's 2026 AI Jobs Barometer, analyzing roughly one billion job advertisements, reports an average wage premium of 62 percent for roles demanding AI skills as of 2026 2. Read that as a reason your existing band will be tested in negotiation, rather than as a number to write into an offer.
There is a leveling trap worth naming. Because the title is new, a candidate's current title tells you very little. Ask what they were allowed to change without approval and how large the catalog was. A person who set publish thresholds across two million titles is doing a different job from someone who curated tags for four hundred, whatever the two business cards say.
On location, the specification, evaluation and policy work is genuinely remote-friendly. What resists distance is the taxonomy argument, which is negotiated across editorial, legal, localization and research and goes badly in a document when the stakes are what the platform is permitted to assert about a title. Distributed teams that run this well bring the owner into the same room as editorial on a regular cadence. Some employers also require on-site work for pre-release or unreleased content, which is a security constraint on the material rather than a preference about desks, and it belongs in the job posting rather than in a surprise at offer stage.
One flag rather than advice. Labels that describe people, assert content warnings, or drive age gating can carry disclosure and record-keeping duties that differ by jurisdiction and are still moving through 2026, and contractual limits from rights holders often bind before any statute does. Check with counsel in your own jurisdiction rather than reasoning from a summary, and expect this role to be the one holding the record of what the system was permitted to say.
Common questions
How do I become a content understanding AI product manager?
Build the artifact the job is made of. Take any catalog you can access, pick fifty items, and run a multimodal model over them to generate tags, descriptions or segments. Then write the taxonomy: definitions precise enough that two people would agree, the boundary cases that break each one, and how you resolve them. Add an evaluation sample weighted toward popular items and consequential labels, score the model against it, and write a publish policy saying which labels may go live unreviewed and at what confidence. That document, plus the errors you found by actually watching the content, is a stronger interview artifact than any course. Metadata operations, localization, accessibility and search relevance are the fastest entry paths.
How is this different from a data product manager or an ML product manager?
Scope of the asserted claim. A data product manager owns pipelines, schemas and freshness. An ML product manager owns a model in a feature. This role owns what the system says about a piece of content, which means the taxonomy, the model that fills it, the threshold at which output publishes without review, and the contracts with downstream consumers such as search and recommendations. The editorial and rights dimension is the part the other two roles usually do not carry. In small organizations one person holds all three; the split appears once the catalog is too large for humans to audit.
Does this role need to be technical?
They need to read and evaluate, not to build. The working requirements are comfort with evaluation results and confusion matrices, enough understanding of multimodal models to know why a label fails on one genre and not another, and the ability to argue with a research team about whether an error is a model problem, a taxonomy problem or a threshold problem. Writing production code is not part of the job. Being able to pull a stratified sample and check labels against the actual content is, and it is the skill most candidates are missing.
Who owns content understanding if nobody has this title?
It is usually split three ways and therefore owned by nobody: research owns the model, metadata operations owns the taxonomy, and a downstream team consumes the field and quietly builds its own assumptions about what it means. The practical test is to ask three people what confidence level lets a genre tag publish without human review. If the answers differ, or nobody knows where the rule is written, the ownership gap is real regardless of what the org chart says.
How large does a catalog have to be before this role is worth hiring?
The threshold is not a title count, it is the point where human review stops being possible and sampling becomes the only quality method available. That usually arrives somewhere in the tens of thousands of items, sooner if items are long-form video where reviewing one takes hours, later if items are short and cheap to check. The other trigger is downstream dependency: once search, recommendations or ad targeting consume a machine-generated field, an error stops being a bad row and becomes a visible product defect.
References
- 1. Netflix careers search, AI Product Manager, Catalog and Content Understanding (Los Gatos) ✓ explore.jobs.netflix.net Discovery evidence that Netflix lists an AI Product Manager role for Catalog and Content Understanding in Los Gatos, pairing the catalog and metadata surface with model-driven understanding of video, work previously split between metadata operations staff and a research team.
- 2. PwC 2026 AI Jobs Barometer pwc.com Supports the macro claim of an average 62 percent wage premium for roles demanding AI skills, across roughly one billion job advertisements, as of 2026.
2 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.