Anyone can make a great anime character now. Making the same one twice is the hard part.
Ten times, in fact. Each one as good as the last. That is what a reference workflow is for. It is the line between a cool sketch and a character you can build a story on.
Spoiler: it holds more than it drops. But the drops are specific, and you should know them going in.
I ran Tsubaki.3, PixAI's newest model, through fourteen reference tests. Seven followed one character through a full production pipeline. Seven went wider on purpose, into creatures, a vehicle, and whole environments. A character reference AI that only holds one face is not much use to most people. Call it a consistent anime character generator or an AI character reference generator. Same job either way: same subject, new scene.
Every test has the exact prompt and a straight read. Let's go.
What held and what broke
- It keeps a character recognizable across new poses, outfits, expressions, and scenes.
- It holds two different characters in one frame without mixing their features.
- It carries creatures, vehicles, and full scenes through big changes too.
- It will not copy a foreign art style onto an existing image. Build that style from scratch.
- In a transform, it keeps the big picture and drops small details you do not restate.
How the reference workflow works
The setup is simple. Generate a base image. Set it as your base or reference from the side panel. Then ask for the same subject somewhere new. The model reads your first image as the source of truth and carries the important parts forward. A reference image AI is only as strong as that first picture, so make it a good one.

To make consistency measurable, I built a test character with clear anchors. Then I named them in every prompt.
an original anime character, a woman in her late twenties with medium-brown
skin, short silver undercut hair with the longer top swept to the left,
heterochromia with one amber eye and one pale grey eye, a thin diagonal scar
across the bridge of her nose, a small gold hoop in her left ear, wearing a
fitted dark teal flight jacket over a grey turtleneck with a worn leather
satchel strap across her chest, calm confident expression, neutral standing
pose, plain light grey background, clean anime illustration, full body

Her anchors: the silver undercut, the amber-and-grey eyes, the nose scar, the gold hoop. Those four are what I tracked.
Part One: one character through a full production
Test 1: the turnaround sheet
Using this exact character, create a professional animation model sheet
showing the same character in four views at the same height on one sheet:
front view, three-quarter view, side profile, and back view. Keep the face,
silver undercut hair, heterochromia, nose scar, gold hoop, flight jacket,
turtleneck, and satchel strap identical across all four views. Neutral A-pose,
plain grey background, clean consistent line work, model-sheet layout.
All four views come out at a consistent scale. You get a usable turnaround. The silver undercut holds across every angle, and the back view shows the shaved side cleanly, which is the hardest angle to get right. The satchel strap carries onto the back too. The hoop stays put.
Two small misses. The nose scar is faint to the point of nearly disappearing at this size. And panels two and three read as two similar near-profiles rather than a clear three-quarter and a true side view, so the second angle teaches you little. The left eye color shifts a touch too.

The reference (left), the four-view turnaround (right).
Test 2: the expression sheet
Using this exact character, create an expression sheet of six head-and-
shoulders portraits in a grid, same character throughout: neutral, wide joyful
laugh, furious anger, quiet grief holding back tears, startled shock, and a
smug half-smile. Keep the face structure, silver undercut, heterochromia, nose
scar, and gold hoop identical in every panel. Only the expression changes.
Plain background, consistent style.

The reference (left), the six-expression grid (right).
This one is built to expose drift. It found a real one.
Her face structure survives the full acting range, including the wide laugh and the shout. That is the harder half to pass. The gold hoop holds everywhere.
But the heterochromia fails in four of the six panels. Both eyes go pale in some, both go amber in others. Only the laugh panel keeps the true split. A single small color anchor is exactly what this model loses track of. If your character has one, plan to fix it by hand.
Test 3: wardrobe swaps
Using this exact character, show the same woman in four full-body outfits side
by side, keeping her face, silver undercut hair, heterochromia, nose scar, and
gold hoop identical throughout: one, casual streetwear; two, a formal evening
suit; three, heavy winter gear with a scarf; four, sci-fi combat armor. Same
height and neutral pose in each, plain background, consistent art style.

The reference (left), the four outfits (right).
The undercut, face shape, and hoop hold across all four outfits. Every figure reads at the same height and build. The armor panel does not secretly bulk her up or slim her down, which is the main failure this test is designed to catch.
The "formal evening suit" came back as a plain navy business suit. Minor prompt miss.
Test 4: film shot coverage
Using this exact character, create a four-panel film shot sequence of the same
character in one location, a rain-soaked neon alley at night, with consistent
wardrobe and one lighting logic throughout: panel one, wide establishing shot
of her standing in the alley; panel two, medium shot from the front; panel
three, tight close-up on her face; panel four, low-angle shot as she looks up.
Keep her face, silver undercut, heterochromia, nose scar, gold hoop, and flight
jacket identical across all shots. Cinematic anime style, consistent neon
lighting.

The reference (left), the four-shot sequence (right).
This is the strongest of the seven. The neon signage, the wet reflective ground, and the alley geometry carry through all four panels. The one-lighting-logic requirement is met, not approximated. The shot progression matches the brief exactly, with no shot-type substitutions.
Best part: the nose scar finally shows up clearly in the close-up. That settles the earlier tests. The model can render the scar. It just loses it whenever the face is small in the frame. The heterochromia reads as two distinct tones in the low-angle panel too. Scale was the limit, not capability.
Test 5: two characters in one frame
First I generated a second, clearly different character.
an original anime character, a lean man in his early thirties with warm beige
skin, chin-length messy black hair with a single blue streak, tired dark eyes,
light stubble, a faded geometric tattoo on the right side of his neck, wearing
a mustard field jacket over a white henley, relaxed skeptical expression,
neutral standing pose, plain light grey background, clean anime illustration,
full body
Then I fed both references into one shot.
Using these two exact characters, place them together in one composed shot:
the silver-haired woman and the black-haired man with the blue streak standing
back to back in a dim workshop, mid-conversation, both looking off in different
directions. Keep each character's face and features exactly as in their
references, the woman's silver undercut, heterochromia, and nose scar, the
man's blue streak, neck tattoo, and stubble. Consistent lighting, cinematic
anime style, both characters fully in frame.
This is the hardest test in the set. The question is whether the two identities stay distinct or bleed together.
They stay distinct. Different hair, different builds, different skin tones, no cross-contamination. You would never mistake one character's features for having drifted onto the other. That is the whole pass here.
Her anchors hold. His blue streak and stubble hold too. His jacket reads more olive than mustard, and the neck tattoo is faint. The "dim workshop" came out more moody than truly dark. The back-to-back staging is right.

Test 5 Two Characters in One Frame
Test 6: cross-format for production and merch
Using this exact character, render the same woman three ways while keeping her
identity identical, her silver undercut, heterochromia, nose scar, and gold
hoop: one, a black-and-white manga panel of her looking over her shoulder with
screentone shading; two, a cute chibi sticker of her giving a thumbs up; three,
a clean flat-color character line-art on white. Keep her recognizable as the
same character across all three formats.

The reference (left), the three formats (right).
This asks the most. The art style itself changes on top of everything else. She still reads as the same person across all three.
The manga panel nails real screentone, not generic shading. The chibi keeps its die-cut border and the exaggerated proportions. The line art is clean and flat as asked. The only real drift is the manga hair reading looser and more windswept than the tidy reference.
Test 7: reference plus pose
Using this exact character, redraw her in this specific pose: mid-stride
running forward, upper body twisted to look back over her shoulder, right arm
reaching behind her. Keep her face, silver undercut, heterochromia, nose scar,
gold hoop, flight jacket, and satchel identical to the reference. Dynamic anime
style, plain background, full body.

The reference (left), the running pose (right).
Drop her into a demanding action pose and she stays herself. Her face is angled toward the camera rather than in pure profile, so this gives one of the cleanest heterochromia confirmations in the series. Amber on one side, a cooler pale tone on the other. The undercut, hoop, jacket, and satchel all hold. The pose is accurate. Only the scar stays faint, consistent with the pattern.
Part Two: beyond characters
The reference workflow does not care whether the subject is a person. The next seven run it on a crew, a scene, a vehicle, a crowd, and two mythic creatures.
Test 8: a graffiti crew, and a style transform
four teenage friends hanging out on a graffiti-covered basketball court at
dusk, a tall boy spinning a ball, a girl with box braids sitting on the fence,
a short kid on a skateboard, a hoodie kid filming on a phone, sketchy urban
anime style, rough ink outlines, marker-textured coloring, visible construction
lines, spray-paint splatter, hand-drawn street-art energy, warm streetlight glow
The base is a good group shot. All four kids interact in one believable space instead of looking pasted together. The spray-paint tags read as part of the scene. The warm dusk glow lands. The one miss is the box braids, which came back as loose curls.
Then I asked the model to re-render it as a 90s anime screencap.
re-render the whole image as an authentic late-1990s TV anime screencap: flat
cel-shaded colors, hand-drawn linework with slight imperfections, chunky
highlights in the hair, muted faded-photograph palette, subtle 35mm film grain,
4:3 broadcast framing, warm nostalgic tone. Make it look like a real frame from
a 90s OVA.

Base graffiti crew (left), the attempted 90s re-render (right).
This is the plainest limitation in the whole review. The 90s look did not take. No flat cel shading, no line imperfections, no film grain, no 4:3 framing. Just a cleaner version of the model's own default style. I ran it twice and got the same result.
Tsubaki.3 will not be talked out of its house rendering by an instruction. Want a specific era or technique? Build it from scratch. Which is exactly what worked later in this batch.
Test 9: a painterly storybook scene
a lone red paper lantern glowing on the branch of an ancient snow-covered tree,
a small round fox curled asleep in the roots below, thick painterly storybook
illustration, visible brush strokes, soft gouache textures, warm lantern light
against cool blue snow, dreamy children's-book atmosphere, no people
This is the best texture result of the set, and it proves the point above. The painterly style was built in from the first pixel, so the brushwork shows up in the bark and the snow-laden branches. The light logic is real too. The lantern's warm glow falls on the snow and bark right beneath it while the rest stays cool blue.
Then the reference move, turning the page to dawn.
Using this exact scene, keep the same tree, the same red lantern, and the same
sleeping fox, but shift it to dawn: the snow now soft pink, the lantern flame
nearly out, the fox stretching awake, first light through the branches. Same
painterly storybook style and brushwork.

Base night scene (left), the dawn shift (right).
The tree, the lantern's position, the pink dawn snow, and the light shafts all carry over. The brushwork stays consistent between the two images. But two specific instructions did not happen. The lantern stayed at full glow instead of burning nearly out. The fox came back in an alert playing pose instead of a sleepy stretch. The model held the scene but ignored the state changes inside it.
Test 10: a hyper-glossy vehicle
a candy-colored futuristic hover-scooter parked on a clean white studio floor,
hyper-glossy hybrid 2D-3D render, smooth cel-shaded surfaces with volumetric 3D
form, glossy reflective paint, soft rim lighting, pastel mint and coral palette,
floating holographic price tag, high-end product-shot aesthetic, toy-like and
premium
The style here is the cleanest technical match in the whole run. The surfaces have cel-shaded banding but real dimensional form. The glossy reflections read as true specular highlights rather than painted-on streaks. The mint-and-coral palette and the product-shot framing are spot on. The only miss is the word "hover." It rests on wheels, grounded.
Then into motion.
Keep this exact hover-scooter identical, same shape, same mint and coral glossy
paint, same details, but place it now mid-drift through a neon night city
street, motion blur, light trails, reflections of signs on its glossy surface,
same hyper-glossy hybrid 3D style.

Base studio shot (left), the neon night drift (right).
Object consistency is perfect. Same shape, same paint, same details. The transform even fixes the earlier miss: tilted and blurred with no ground contact, it now reads as airborne. Neon signs smear across its glossy panels exactly as asked. One unrequested extra: a blurred human silhouette in the corner that nobody asked for.
Test 11: a desaturated crowd scene
a crowded rain-soaked train platform at night, dozens of anime commuters with
umbrellas, one still figure standing motionless in the flowing crowd,
desaturated cinematic anime style, muted grey-blue palette, single warm sodium
lamp as the only color, heavy rain, shallow depth of field, film-still framing,
melancholic mood
This is the model's best literal read of a color-and-mood brief. The whole frame is cool grey-blue except the one warm lamp, exactly as specified. The crowd is dense and real. The depth of field falls off properly. The still figure stands out because her umbrella is down while everyone else's is up. Stillness shown through an action, not just position.
Then the camera move.
Using this exact scene, keep the same crowd, the same still figure, the same
rain and desaturated palette, but change to a low tracking shot from behind the
still figure looking out at the departing train, wet reflections on the
platform, same cinematic desaturated grade.

Base crowd shot (left), the tracking shot from behind (right).
The camera reposition works. The wet platform reflections are excellent. But this one drops the "keep the same" parts. The crowd thins from dozens of people to about four. The warm-lamp accent that made the original distinctive is gone, leaving everything uniformly cool. The scene held its mood but lost the specific elements that defined it.
Test 12: a neo-cyberpunk market
a dense neon night market in a neo-cyberpunk megacity, narrow alley packed with
holographic food stalls, steam and smoke, hanging cables and glowing signs in
Japanese and Korean, mixed crowd of cybernetic vendors and shoppers, reflective
wet ground, teal and magenta neon, chromatic aberration, blade-runner anime
atmosphere, deep perspective
The environment is excellent. The dense alley, the hanging cables, the steam, the mixed signage, the wet neon reflections, and the deep perspective all land. The market feels packed without turning to chaos. The one weak spot is the people. The crowd works at a scene level, but individual faces are soft and a little distorted. The model renders the place far better than the people filling it.
Then a shift in hour and weather.
Keep this exact market identical, same stalls, same signs, same layout and
perspective, but shift it to a quiet pre-dawn: most signs off, one lit noodle
stall, a lone street cleaner, fog rolling through, cool blue palette with a
single warm stall glow, same neo-cyberpunk style.

Base night market (left), the pre-dawn version (right).
The pre-dawn mood is strong. Cool blue, heavy fog, one warm noodle stall, a lone cleaner. But "keep this exact market" only half holds. The layout and stalls carry over while the detailed signs turn into blank panels. Same lesson as the crowd scene. The model keeps the shape of a place but not its fine print.
Test 13: a retro 90s mecha
a battered military mecha kneeling in a wheat field at golden hour, one arm
damaged and smoking, birds scattering, distant mountains, authentic late-1990s
OVA anime style, flat cel-shaded colors, hand-painted background, heavy line
weight, muted faded palette, 35mm film grain, 4:3 framing, nostalgic
broadcast-still look
Built from scratch, the retro look lands far better than it did as a transform in Test 8. The heavy outlines, the muted military palette, and the film-like texture give a convincing retro feel, though it leans modern-retro rather than a true 90s cel frame. The mecha, wheat field, mountains, and golden hour are all there. A miss: the damage reads more like burning debris than subtle smoke.
Then the hero shot.
Keep this exact mecha identical, same battle damage, same design and faded 90s
palette, but pull back to a dramatic wide low-angle hero shot as it slowly rises
to standing against a burning orange sunset sky, dust and embers, same 90s
cel-animation OVA style and film grain.

IBase kneeling mecha (left), the rising hero shot (right).
A strong transform. The exact mecha design, palette, and battle damage carry straight over. The move from kneeling to a low-angle standing shot makes it imposing. The sunset, embers, and smoke land well. The rendering is a touch more refined than a real 90s frame, but the identity and the drama both hold.
Test 14: an ukiyo-e koi-dragon
a great white koi transforming into a dragon as it leaps up a roaring waterfall,
traditional Japanese ukiyo-e woodblock style fused with modern anime color, bold
flat outlines, wave patterns like Hokusai, gold-leaf accents, indigo and
vermilion palette, visible paper texture, mythic and elegant, no people
This is the strongest single result in the whole set. The woodblock influence is unmistakable. Flat color, decorative wave patterns, gold cloud accents, paper-like texture. The creature reads clearly as a koi mid-transformation, scales and fins alongside a dragon head and horns. The palette and composition are excellent. The only nitpick: it looks more like a dragon emerging than the exact instant of change.
Then the final transformation.
Using this exact koi-dragon, keep the same design, palette, and woodblock-anime
style, but show the final moment of its transformation at the top of the falls,
now a full serpentine dragon coiling into storm clouds, lightning, same ukiyo-e
woodblock fusion and gold-leaf accents.

Base koi-dragon at the falls (left), the final storm-cloud form (right).
The reference move is clean. The same design, mane, horns, and palette carry over. The subject coils up into storm clouds with lightning drawing the eye, a strong vertical composition. The one honest limit: the creature keeps its koi head and fins, so it lands between a koi-dragon and a fully serpentine one rather than completing the change. And the rendering drifts a little more cinematic than strict woodblock.
The pattern across all fourteen
Across all fourteen, the model is good at one thing in particular. Give it a subject and ask for it again somewhere new, and it holds. A new pose. A new angle. A different time of day. A change in weather. It kept two separate characters in one frame without mixing their faces. That is the main thing you want from a consistent character AI, and it does it.
The weak spots are clear too. It will not copy a different art style onto an image you already made. It falls back to its own look every time. It drops fine facial detail when the face is small, then draws it correctly once you move in close. And in a transform it keeps the overall scene but loses the small things you do not mention again. A sign goes blank. A crowd thins out. A lamp you asked to dim stays lit. Name those things and they stay.
How to get consistent results
The whole review turns into a short set of habits.
- Start with a clean, well-lit reference facing forward. Every later image inherits it.
- Name the details that define your character, the exact hair, eyes, and signature items.
- Check consistency in close and medium shots, where fine detail shows up.
- For a specific art style, build it from scratch instead of converting an existing image.
- When you change one thing, list everything you want kept, not just the thing you are changing.
- Use it for more than faces. It holds objects, creatures, and scenes just as well.
PixAI's Reference Pro guide covers the reference workflow in full. And once you have a character you want to reuse, training a character LoRA on PixAI turns it into a model you can call up any time, so you are not re-uploading a reference for every image.
Your next step
For keeping a character consistent across poses, outfits, expressions, scenes, and shots with a second character, Tsubaki.3 does the job. It works as an OC character generator and a scene tool in the same breath. Its limits are small and easy to work around once you know them.
The way to know if it fits your work is to test it. Take your own character, run it through a few scenes on PixAI, and see what holds. You will know in about three images.
















