Keeping AI characters consistent across scenes is one of the most common pain points for comic creators, children's book authors, and social media teams working with AI art. The core problem: generate a perfect face, then try to recreate it in a different pose or setting and get a completely different person. The tools for solving this have improved dramatically in 2026 — Midjourney v7 introduced better character reference handling, FLUX models have stronger identity preservation than any earlier version, and Stable Diffusion's ComfyUI ecosystem now has purpose-built consistency nodes. This guide covers the specific workflows that work today across each major platform.
Last updated: August 3, 2026
Midjourney v7: --cref and Improved Character Consistency
Midjourney's --cref (character reference) parameter remains the simplest consistency solution available. The workflow: generate one strong portrait of your character, then use that image URL as a reference for all subsequent generations.
Basic syntax: Woman in a coffee shop, casual outfit --cref [image-url] --cw 100
The --cw (character weight) parameter controls how strictly the new image matches the reference face. --cw 100 preserves exact facial features; --cw 75 allows slight variations while keeping the character recognizable; --cw 50 captures the general vibe without strict face matching. Midjourney v7 (released early 2026) made two meaningful improvements: better handling of non-frontal reference poses, and stronger hair and eye color preservation that was inconsistent in v6.1.
Combine --cref with --sref (style reference) to maintain both character identity and the overall artistic style across a series — especially useful for illustrated children's books where you need consistent painting style alongside consistent faces.
One important caveat: --cref works best with photorealistic and semi-realistic styles. For anime or heavily stylized characters, the results are less consistent — use the Stable Diffusion LoRA approach instead.
Browse Midjourney portrait prompts on PromptSpace for starting points that generate strong hero portraits suitable as character references.
Stable Diffusion: LoRA, IP-Adapter, and InstantID
Stable Diffusion offers the most options for character consistency, each with different trade-offs between quality, training time, and ease of use.
LoRA Training (Gold Standard)
Train a custom LoRA on 10–20 high-quality reference images of your character using Kohya_ss or CivitAI's online trainer. Training takes 30–60 minutes on a mid-range GPU and produces a small file you can reuse indefinitely. The results are remarkably consistent across different poses, outfits, and lighting because the LoRA teaches the model your character's specific facial structure. Trigger it in your prompt with the LoRA activation syntax and the character appears faithfully every time.
Best training images: sharp focus, varied lighting conditions, multiple angles (front, three-quarter, profile), clean backgrounds, consistent art style across all reference images. 20 images beats 10 for most characters; more than 30 shows diminishing returns.
IP-Adapter (No Training Required)
Load a single reference face image and IP-Adapter conditions the generation to match those facial features — no training required. IP-Adapter Face focuses exclusively on facial features and gets you 80% of a LoRA's quality with 0% of the setup time. For professional work like comics, brand mascots, or visual novels, invest the time in a LoRA. For quick projects, IP-Adapter is the faster path.
InstantID and ReActor
Both take a face-swapping approach: generate any image, then swap in the face from your reference. InstantID produces more natural-looking results. ReActor is faster but can look slightly composited on close inspection. Both require a clear, well-lit reference face — blurry or dark references produce artifacts at the face boundary.
FLUX: Reference Image Workflows in 2026
FLUX's photorealistic foundation makes it particularly strong for character consistency — because FLUX generates convincing faces by default, maintained features look naturally integrated rather than pasted. Main tools for FLUX character consistency:
- FLUX PuLID and FLUX IP-Adapter — available through platforms like Replicate and fal.ai. Provide a reference face image alongside your text prompt. FLUX handles profile views, three-quarter angles, and dramatic lighting changes while preserving core identity features better than earlier models.
- FLUX Controlnet with face depth maps — for controlling both pose and identity simultaneously. More technical to set up but gives precise control over body position and facial direction.
- FLUX LoRA training — increasingly popular as FLUX has become the default photorealistic model. Training tools for FLUX LoRAs have matured significantly in 2026.
For best results: use a clear, well-lit, forward-facing reference photo. Describe the scene, outfit, and setting in your text prompt — let the reference image handle facial identity. Find starting-point prompts at 50 FLUX Prompts for Portrait Photography.
DALL-E 3 / ChatGPT: Conversational Consistency
DALL-E 3 via ChatGPT lacks dedicated reference image parameters, but achieves reasonable consistency through conversation context and detailed character descriptions.
- Write a comprehensive character sheet with every physical feature: hair color, length, and style; eye color and shape; skin tone; facial structure; age; distinguishing marks.
- Use this exact description as a prefix for every generation within the same ChatGPT conversation thread.
- Build on each previous image: "Same character as above, but now she's in a train station at night, wearing a long coat."
Results aren't pixel-perfect, but the character stays recognizable across scenes within a single conversation. Works best for illustration-style images rather than photorealistic portraits. Use the DALL-E 3 character design collection as a starting point for generating a strong initial character sheet image.
The Multi-Tool Professional Workflow
The most reliable approach for professional character work combines multiple techniques:
- Generate a hero portrait. Front-facing, clean lighting, no distracting background. Generate several variations and select the best.
- Write an exhaustive text description. Document every physical feature specifically enough that the description alone produces consistent outputs even without an image reference.
- Build a multi-angle reference sheet. Generate front, three-quarter, and profile views. Models maintain identity more reliably when they've "seen" the character from multiple angles.
- Use text description and image reference together. The combination produces more consistent results than either alone.
- Generate in batches of 4 and select the best. Always go back to your hero portrait as the reference — never use a recent (potentially drifted) generation as the new reference.
Common Mistakes That Break Consistency
Changing too many variables at once. Move scene and outfit but keep lighting conditions roughly consistent until the character is strongly established. Jumping from soft daylight to dramatic neon in the second generation creates enough style tension that the model deprioritizes face matching.
Extreme style jumps. Going from photorealistic to anime within the same series produces completely different faces. Stay within one style family.
Using recent generations as the reference. If image 7 drifted slightly, using it as the reference for image 8 compounds the drift. Always reference the original hero portrait.
Describing the face in every prompt when using reference images. Let the reference handle facial identity. Focus your text on scene, outfit, lighting, and composition — conflicting text and image instructions reduce consistency.
Getting Started: Find Your First Hero Portrait
Browse PromptSpace's Midjourney portrait prompts and DALL-E 3 character design collection for templates that produce strong, reference-quality faces. Start with a template, customize the physical description, generate 4–8 variations, pick the best — and that becomes your reference image for everything that follows. For a deeper look at AI portrait photography specifically, the AI headshot prompts guide covers lighting and composition principles that apply directly to character hero portraits.
Frequently Asked Questions
Inconsistently. The --cref parameter was designed primarily for photorealistic and semi-realistic styles. For anime characters, a Stable Diffusion LoRA trained on your character design produces more reliable results. If you're committed to Midjourney for anime work, use --cref at --cw 75 (not 100) and generate larger batches — consistency varies significantly by generation.
10–20 high-quality images is the practical sweet spot. Fewer than 10 tends to produce a LoRA that only reproduces the reference exactly without pose or lighting flexibility. More than 30 shows diminishing returns. The 15-image set — front, three-quarter, profile, various expressions, varied lighting — covers the cases you'll typically need.
Technically yes, but there are legal and ethical considerations. Using photos of public figures without permission may violate their rights of publicity depending on your jurisdiction and use case. Photos of private individuals require explicit consent. The safest approach: use an AI-generated hero portrait as your reference image — no rights issues and complete control over the character's appearance.
DALL-E 3 through the free tier of ChatGPT with a detailed text description is the most accessible starting point. Stable Diffusion (runs locally, completely free) with IP-Adapter is the most capable free option. Midjourney requires a paid subscription. FLUX-based tools are available through free-tier API access on platforms like Replicate with usage limits.
Three things prevent drift: (1) always reference the original hero portrait, not a recent generation; (2) generate in batches of 4 and select the most on-model result rather than accepting the first output; (3) build a multi-angle reference sheet early and use the most relevant angle depending on the scene's composition. If drift has already occurred, regenerate the drifted images from scratch using the original hero portrait rather than trying to correct them.












