Inside Zip AI’s character bibles, Higgsfield identity tests, and the search for a convincing speaking performance
A character can look convincing until the moment you ask for the next shot. The first portrait has the right face, the right clothes, the right expression. Turn the camera sideways and the jaw changes. Move into a brighter room and the hair changes. Add dialogue and the mouth seems to belong to a different performance. You have a collection of attractive images, but you still do not have someone an audience can follow from one video to the next.
That is the practical problem behind Zip AI’s virtual staff. Our characters are fictional, AI-generated people used to explain real ideas about marketing, media, engineering, operations, finance, and logistics. They need distinct identities and recognizable personalities. They also need to remain those characters when a different tool, a different shot, or a different editor enters the process.
This article follows the work behind that goal: the character bible, the reference library, the identity tests, the voice decisions, and the difficult transition from still image to speaking performance. Harry, our logistics character, provides the most complete documented visual example. The wider staff shows why keeping production records straight is part of the creative work. The next speaking test may use HeyGen, but that is a test we still need to run, not a success we can claim in advance.

Begin with a character bible that can actually be used
A useful character bible answers two questions at once. Who is this person within the fictional world, and what must the production team preserve to make the next appearance believable? The first question gives a writer something to write. The second gives an image or video operator something to check. Neglect either one and the result feels unfinished.
A biography alone is not enough. A paragraph describing a confident, approachable expert gives little guidance when a scene requires that expert to disagree with someone, admit a mistake, or explain a confusing tool. A technical asset sheet alone is not enough either. You can attach the right face and use the right voice while writing dialogue that could belong to any member of the cast.
Our working structure gives each character a role, an audience promise, a visual reference, a speaking style, a range of emotional responses, a wardrobe policy, and a production record. It also gives uncertainty a visible home. An unconfirmed voice setting stays unconfirmed. A proposed habit stays a proposal. An old draft does not become approved history merely because someone copied it into a newer document.
Consider a finance character. ‘She is analytical’ describes a category. A direction such as ‘she asks for the total cost before responding to the promised benefit’ gives a writer an action and an actor a sequence. It can shape a line, a pause, or a reaction shot. It still leaves room for warmth and surprise. She does not have to reject every purchase to remain recognizably interested in financial tradeoffs.
The same test helps separate the other roles. The engineer wants to understand what failed and whether the fix can be reproduced. The operations lead wants to know who owns the next step and whether another person can repeat it. They can both be methodical without having the same temperament or speaking in the same rhythm. This distinction belongs in the bible before the dialogue is generated.
Give each character a way of seeing the same problem
For the current staff work, the six leads are Freya Lind, Valentina Cruz, Kenji Mori, Marcus Hale, Eleanor Vance, and Harry. Their department labels are useful, but the labels cannot do all the characterization. Six people who calmly explain their specialty will soon sound like one presenter wearing six outfits.
A stronger approach is to give them different questions. Freya, in marketing and branding, asks who the message is for. Valentina, the AI-generated media operator, asks how an idea becomes a finished piece of work. Kenji asks how the system behaves. Marcus asks how people will operate it. Eleanor asks what the decision costs and what it displaces. Harry asks whether the operation still works when real-world conditions change.
These are writing lenses, not invented proof that the characters have lived particular lives. We do not need to manufacture a childhood injury, a failed marriage, or a dramatic employment history to make a fictional presenter specific. A backstory can be useful, but it should be deliberately written, reviewed, and connected to something the audience can observe. Otherwise it adds obligations without improving a scene.
The bibles also need room for behavior that is not a job description. A character might become unexpectedly animated when explaining a small discovery. Another might use a dry joke to make a nervous beginner comfortable. These details create texture when used selectively. Repeating the same pen click, eyebrow raise, or catchphrase in every appearance quickly turns a person into a mechanism.
Relationships need the same care. Professional overlap can suggest good scenes, but friendship, rivalry, reporting lines, and shared history should be established deliberately. In this project, some early ensemble proposals went beyond the existing canon. They remain development ideas. A writer can propose an interesting disagreement without retroactively declaring that two characters have always been rivals.
An established face is not an invitation to cast again
One of the clearest failures in this project happened before animation. Three staff characters already had reference photographs and production history. Their reusable identity setup was incomplete or inconsistently documented. That status was misread as an absence of established faces, and new casting candidates were made for people who had already been cast.
The production handoff records twelve rejected images and 36 credits spent on that mistaken round. We are reporting the recorded figure, not a newly reconciled wallet statement. None of those rejected candidates belongs in the characters’ reference packages. Their value to this article is the lesson about source selection, not the opportunity to recycle a face that looks appealing.
‘Not trained’ and ‘not cast’ describe different situations. Casting chooses a character’s appearance. A reusable identity asset helps preserve an appearance that has already been chosen. If an existing character needs better side views or a more reliable reference package, the assignment is to extend and validate the established identity. It is not to choose a different person.
The practical check is simple: open the actual reference images before writing a new human-likeness prompt. Look at the published or approved material. Read the decision record that made a particular image authoritative. A folder name and a summary are helpful directions to evidence; they are not substitutes for looking at it.
This matters especially when work is spread across old computers, copied folders, and multiple assistants. A directory can contain both the original actor and later experiments. ‘Final’ can mean a technically exported file rather than a creatively approved one. ‘Accepted’ can refer to a shot without making that shot the new source of identity. The bible should identify the exact approved anchor and the purpose of each supporting asset.
Treat the reference library as evidence
During the audit, filenames hid useful material. An image inside Freya’s folder had a generated identifier rather than ‘side profile’ in its name. Opening it revealed a side view. Kenji’s original named folder lacked a profile, but an archived identity-rebuild package contained profile references. Finding those files did not automatically prove their approval, but it did change the problem from ‘create a missing view’ to ‘verify the existing view and its history.’
The opposite problem also appeared: one identical image was filed under two different characters. Comparing file hashes confirmed that the copies contained the same bytes. That proves duplication. It does not prove which character the image was intended to represent. The safe production response is to withhold that ambiguous image from identity inputs until its source assignment is established.
A reference package should therefore carry more than pictures. Each image needs a stable name or identifier, its source, its intended role, and its status. Is it the identity anchor, an approved wardrobe reference, a full-body check, a scene-specific composition, or a rejected diagnostic? A small amount of clear labeling prevents a great deal of later confusion.
Record visual observations narrowly. Hair length may be clear in a front portrait. Eye color may not be reliable in a small, strongly graded image. A ring may be visible, but it is an accessory rather than facial anatomy. An uneven eyelid in one frame may be an expression or generation error. It should not become a permanent feature because it happened to appear once.
For each proposed invariant, keep the image reference, the visible feature, the viewing conditions, and whether other approved images support it. Approval and confidence are separate fields. You can be highly confident that an image contains a scar while having no authority to make that scar part of the character. That distinction would have prevented several of our early mistakes.
Use Higgsfield identity tools for the job they actually perform
Higgsfield’s current Soul ID guide describes training a reusable identity from a varied set of portraits, with 20 or more images and support for up to 80. It also says invented characters can be trained from generated portraits. The important practical requirement is that the images consistently represent one face. Training cannot turn a mixed collection of different-looking people into a well-defined character by itself.
The same guide distinguishes a trained Soul identity from supplying a reference image for a single generation, and explains how trained characters can be used through Elements in supported workflows. Keep those concepts separate in your records. A portrait, a trained identity, a reusable Element, and a generated scene are related assets, but they are not interchangeable.
Source: Higgsfield: Create and use a Soul ID character
For an established Zip AI character, our first task is to recover the asset that already represents that character and verify its sources. We should not infer live account state from an old document that says ‘trained’ or ‘not submitted.’ We should also avoid treating a generic model name in a historical plan as a current instruction. The interface, supported inputs, and account availability need checking at the point of production.
When preparing a reference set, use clean views that show useful information. A clear face, a useful angle, and an unobstructed silhouette help a reviewer compare results. A dramatic color grade or an elaborate costume may be perfect for a finished scene while being poor evidence for eye color, hairline, or body proportions. Keep scene styling out of the identity decision whenever it obscures the thing being judged.

Harry shows what a controlled identity test can reveal
Harry’s package begins with a role: a fictional logistics operator who can explain aviation, maritime operations, and complex physical movement while keeping human judgment central. His casting brief called for a calm, experienced presence. The production record identifies Candidate A as Rodney’s selection. From that point, the process should preserve that actor rather than keep searching for a more attractive replacement.
The historical proof set tests front and angled views, profiles, full-body framing, seated and walking poses, stronger side lighting, and a wardrobe change. Its record marks ten views as passed. The point of showing that sheet is not to promise perfect identity consistency. It is to make the evaluation visible: a reviewer can compare the face and proportions across conditions instead of approving one flattering portrait in isolation.

A useful review begins with the anchor image beside the new result. Compare the overall face before getting absorbed in texture. Look at the hairline, jaw, nose, apparent age, and proportions. Then examine the eyes, teeth, hands, and areas where clothing meets the body. A result can be beautiful and still fail because it reads as someone else.
Next, change the environment. Harry’s library includes command-center, helicopter, and yacht imagery. These are fictional production scenes, not evidence that the character operates real vehicles or holds real credentials. Their production purpose is to test whether the actor remains recognizable while setting, wardrobe, and lighting change.
One command-center attempt was rejected for malformed interface labels. The corrected version changed the scene treatment. That is a useful failure to retain because it shows that identity is only one dimension of acceptance. A stable face does not excuse unreadable text, misleading branding, or a scene whose visual logic does not work.

Realism is not a checklist of imperfections
There is a tempting shortcut in AI portrait work: if the face looks too polished, add a scar, a wrinkle, an asymmetry, or a tired eye. Specific instructions can affect a result, but that does not make the requested feature right for the character. The goal is a believable, continuous identity, not a quota of physical irregularities.
In the rejected development round, physical details were invented for established characters. The instructions included features that contradicted the references or the producer’s preferences. Some of those requested details reportedly rendered, but a successful rendering of the wrong detail remains a failed character result. Instruction following and creative correctness are different tests.
That round also led to a proposed rule: specifying an imperfection works better than forbidding airbrushing. We are not presenting that as a proven finding. The experiment lacked a matched control and did not isolate model choice, lighting, reference strength, or random variation. It supports a narrower observation about those particular outputs. A stronger general claim would need a properly controlled comparison.
Our practical approach is to preserve visible, approved details and ask for a coherent photographic treatment. Skin, lighting, expression, and anatomy should work together. If a face changes whenever the lighting changes, adding more adjectives about pores will not resolve the underlying continuity problem. Return to the reference chain and the identity test.
A voice is part of the character record
The audience may recognize a voice before it sees a face. A production bible should therefore preserve the selected voice, the provider and model when known, the preset or clone identifier, the approved audio sample, and the delivery settings actually used. ‘Warm male voice’ is a casting description. It is not enough information to reproduce an established performance.
Harry’s voice record identifies Cillian as the selected preset. Other characters have different levels of documentation. Some have a saved performance with incomplete provider metadata. Some have an explicit unassigned status. These cases need different treatment: preserve an existing performance where possible, recover missing settings, and audition an unassigned voice deliberately. Do not replace a voice merely because a convenient default is available.
The archive also showed why dates and decision history matter. An earlier newsroom voice sheet conflicted with later production. A subsequent written lock recorded Bella with Celine and Emily with Faye, and the saved production audio supported those later assignments. The useful lesson is to find the actual later decision rather than ask the producer to repeat a choice or assume the oldest document with ‘locked’ in its title must win forever.
Delivery belongs beside the identifier. The same preset can sound unlike itself if one clip is rushed, another is pitched differently, and a third uses exaggerated emotional direction. Keep a short reference performance that demonstrates the normal rhythm, pronunciation, and emotional range. A character should be able to sound pleased or concerned without becoming a different speaker.
Write for speaking before generating audio. Shorten clauses that are difficult to say naturally. Mark a meaningful pause where the character changes thought. Confirm how names and the brand should be pronounced. Listen to the entire approved recording before animating it. Changing the audio afterward can invalidate an otherwise usable synchronized performance.
The mouth exposes a problem that the portrait can hide
A still image only has to hold for one moment. A talking shot must preserve identity while the jaw, lips, cheeks, teeth, eyes, and head move together over time. The performance also has to follow the audio. That makes speech a separate production problem, even when the same platform offers both image and video generation.
Harry’s recovery documentation describes an earlier assembly that placed voice over footage without a satisfactory synchronized performance. That assembly was rejected. The record later describes a dedicated speaking pass and a rebuilt introduction. We can use that history to explain the workflow, but a documentary must not substitute a written ‘passed’ label for showing the actual motion it is discussing.
Start a talking test with one visible speaker, a clear mouth, restrained movement, and approved audio. Keep the initial camera simple. A sweeping move, a turning head, a hand crossing the face, and another person entering the shot make it harder to diagnose the reason for failure. The first test should tell you whether the face and speech work together.
Watch at normal speed with sound first. Does the performance feel synchronized, or does the mouth lag behind the voice? Then inspect the difficult moments: closed-lip sounds, broad vowels, pauses, and transitions into or out of a smile. Check whether teeth change shape, the jaw loses its structure, or a closed mouth keeps moving during silence. Return to normal speed before deciding; a freeze-frame alone can make ordinary motion look stranger than it feels.
Expressions deserve their own review. A character can articulate words acceptably while looking emotionally unrelated to them. Constant smiling during a serious explanation can be as distracting as a timing error. The bible’s performance direction should connect expression to thought: the character considers, decides, explains, or reacts. Motion for its own sake is not the same as an intelligible performance.
Why our next speaking test may use HeyGen
Our working hypothesis is that a dedicated presenter workflow may produce a more useful speaking performance for these characters. HeyGen’s current Single Scene documentation describes creating a Photo Avatar, supplying a script or uploaded audio, and using Presenter Mode for direct-to-camera delivery. That makes it a relevant candidate for a controlled comparison using the character and recording we already want to preserve.
There is an important distinction inside HeyGen itself. Its documentation separates Presenter Mode from Cinematic Mode, which it describes as powered by Seedance. Simply moving the job to a different website would not necessarily test the different kind of speaking engine we are interested in. We need to record the actual mode and engine used, not just the product name.
Source: HeyGen: Animate a Photo Avatar with Single Scene
The comparison should use the same approved portrait and audio wherever the tools support them. Keep framing, duration, and motion direction comparable. Judge speech timing, mouth anatomy, identity retention, expression, and the amount of cleanup required. Record the cost of the usable result, including rejected attempts. A more expensive first render can be cheaper than several unusable retries, but that is something to measure rather than assume.
No HeyGen comparison result is claimed in this article. We have a reason to test it and a way to judge it. Until that test is complete, the honest conclusion is that Higgsfield identity work and a separate speaking workflow may complement each other. We will choose the route that preserves the character and produces an acceptable performance, rather than force every stage through one tool.
Make one controlled change when a test fails
Unstructured retries are difficult to learn from. If the next attempt changes the face reference, the audio, the framing, the model, and the motion prompt, you may get a better result without knowing which change helped. You may also lose something that was already working. The production record should make the difference between attempts explicit.
For example, if a microphone obscures the lower lip, change the composition before trying to solve the problem with a longer performance prompt. If the face drifts even in a simple close-up, inspect the identity inputs. If the voice sounds wrong, solve the audio before paying for another animation. If only a difficult camera move fails, keep the accepted speaking setup and simplify the shot.
Give each test a specific acceptance question. ‘Does this look good?’ is too broad to guide an expensive iteration. ‘Does this approved face remain recognizable through a twelve-second explanation with one small head movement?’ gives a reviewer something concrete to answer. Additional ambition can follow after the basic performance passes.
A rejected output should remain labeled as rejected. Retaining it for a documentary or technical comparison does not make it an alternate master or an identity source. Keep its prompt, its result, and the reason for rejection together. Otherwise the next person searching the folder may mistake a polished-looking failure for a useful reference.
Capture the process before it disappears
If you want to teach this workflow, documentation cannot wait until the edit. Capture the setup before submitting: the selected reference, the visible model or mode, the prompt, and the relevant controls. Then capture the result and the comparison that led to the decision. Screenshots of the completed library are useful, but they do not reconstruct every choice that produced it.
For stills, save the original output and an annotated review copy. For speaking tests, keep the original video and audio as well as the review notes. A contact sheet can reveal changes in a face, but it cannot prove synchronization. A screen recording can show the workflow, but it should not become the delivery master when the source video is available.
The Harry archive contains a concrete reminder of that last point. Its recovery script describes a browser-recorded assembly that appeared finished but had been reduced to one frame per second by background-tab throttling. The documented recovery rebuilt the edit from source media. That is a technical failure a still screenshot could never expose. The exported file needs its own playback and stream checks.
Keep historical evidence distinct from a recreated tutorial. If an original setup was not captured, demonstrate the current interface and label the demonstration as a reconstruction. Do not arrange a screenshot so it appears to document an original job it never witnessed. The article becomes more useful when readers can tell what happened, what was recovered afterward, and what remains to be tested.
Before publishing screenshots, crop unrelated tabs, account information, private links, and other projects. Show the controls that explain the decision. Preserve the unaltered originals privately so the production record remains intact. The public figure and the internal evidence file serve different purposes.
The handoff that makes the next episode easier
At the end of a character build, another operator should be able to continue without asking who the character is, which voice to use, or which folder contains the real reference. That is the practical standard for the bible. It should point directly to the identity anchor, supporting views, approved voice sample, known settings, wardrobe rules, writing guidance, and representative accepted performances.
Include the unresolved items, but make them specific. ‘Voice incomplete’ is less useful than ‘approved audio exists; provider identifier not yet recovered.’ ‘Needs references’ is less useful than ‘front and full-body views verified; profile derivative located but approval unconfirmed.’ Precise gaps are easier to resolve, and they discourage a new operator from inventing a replacement.
A short production example belongs in every bible. Give the character a few lines that demonstrate the intended cadence, one simple gesture, and a clear emotional purpose. Mark new writing as proposed until approved. Once a sample performance is accepted, preserve it alongside the instructions. Examples help future writers interpret the rules without turning every appearance into an imitation of the same shot.
For the larger cast, use a shared structure with separate character records. Common fields make the package easy to navigate. Separate references keep the identities from bleeding into one another. The individual character should remain distinct even when a costume, setting, or technical workflow is shared.
What we would do first on the next build
We would begin by checking whether the character already exists. We would establish the exact visual anchor, then write a compact, performable character description. We would prepare and review the reference package, confirm the voice, and prove a simple speaking shot before committing to elaborate motion. Each stage would leave behind evidence that makes the next stage easier to judge.
We would also decide what we are making. A documentary about the build and a commercial starring the finished character need different footage and different editorial choices. They can share assets, but neither should disappear into the other. A compelling process film shows decisions and consequences. A character introduction gives the audience a reason to remember the person.
That is the standard we are applying to the Zip AI staff. The face matters. So do the words, the voice, the pauses, and the consistency of the next appearance. The tools will continue to change. A clear character bible and an honest record of what worked give the production something stable to build on.
If you are starting your own recurring character, begin with one you can define, one reference you can defend, and one short performance you can evaluate. Keep the result that works. Record why it works. Then make the next scene.
More from Zip AI
Building Harry: the original character production story
How to create a consistent character in Higgsfield
