The first outfit swap I liked was still wrong. The jacket was convincing, the light matched, and then I looked at the collar. The model's face was fine, but the garment had become a vague cousin of the original piece. That test taught me the most useful rule for an AI model outfit swap. The person is an identity problem, but the garment is a product-accuracy problem. They need different evidence.
Since then, I run every swap as a small fitting session inside imgmov. I upload a consented model photo, add clean garment references, keep the pose and background locked, and only then let the model change clothing. The workflow below is the version I now reuse when I need a set of lookbook-style images without reshooting every piece.
Prepare the evidence first
A good source model photo answers most of the workflow before generation starts. I look for a full-body adult subject, plain studio light, a pose I can repeat, and clothing that does not hide the silhouette I need to replace. Hands matter more than I expected. If the fingers disappear into a pocket or cross the garment seam, the model has to invent fabric around them.
My pre-flight list is short:
- I own the model photo or have written permission to use the person's likeness.
- The subject is an adult and the styling is safe for a commercial context.
- The photo is sharp, at least 1024 pixels on the short side, and has no watermark.
- The face, hairline, neck, hands, feet, and body proportions are visible.
- The pose leaves enough room around the garment to show its real shape.
The clothing reference is just as important. A front view is rarely enough for a tailored jacket, wrap dress, pleated skirt, or anything with hardware. I prefer a front reference plus a back or detail shot when the garment has closures, embroidery, a distinctive neckline, or a special hem. I crop away mannequins, hangers, floors, and price tags when possible. The model should receive the garment, not a room full of retail furniture.
Define the two roles
My prompt starts by naming the references. Otherwise the model may average the two images into a new person wearing an imagined outfit.
Reference 1 is identity. Keep the same adult model, face, hairstyle, skin tone, body proportions, pose, camera angle, framing, and background.
Reference 2 is garment only. Replace the current outfit with the exact garment in Reference 2: [color, material, neckline, sleeves, closure, length, fit].
Preserve natural fabric folds, seams, shadow direction, and skin texture. Do not beautify the face, reshape the body, add a different person, add text, or add a watermark.
That phrase, "garment only," does a lot of work. It separates the face I want to preserve from the piece I want to move. If I am swapping a top and keeping trousers, I also state the trousers explicitly. A model will otherwise decide on its own.
For a first test, I keep everything constant except one variable. The same model photo, same camera framing, same background, same prompt structure. I change only the garment reference. If I change pose, scene, fabric, and accessory at once, I cannot tell which instruction caused the failure.
Choose the model by failure mode
On imgmov, I no longer ask which model is best in general. I ask which failure I can afford to fix.
| Model | When I use it | Why |
|---|---|---|
| Agnes Image 2.1 Flash | Guest rehearsal only | It supports image-to-image, but guest previews are watermarked, limited to three per day, and not for commercial delivery. |
| Seedream 5.0 Lite | Primary photoreal outfit swap | Supports multiple image references and costs 5 credits per image. |
| GPT Image 2 | Complex layering, readable labels, or graphic garment details | Supports multiple image references and costs 2-80 credits by quality and resolution. |
| SDXL Inpainting | One damaged area in an otherwise approved image | Costs 3 credits per image and works with a mask. |
| Nano-Banana 2 | Text-only concept boards, not the exact swap | It is useful for exploring wardrobe mood before I spend credits on production references. |
Agnes is my rehearsal layer when I am logged in as a guest. It can tell me whether the prompt roles are clear, but I never ship a watermarked preview. For the production route, Seedream 5.0 Lite is usually my first choice because it handles multiple references and keeps photographic context stable. GPT Image 2 earns its credits when a look includes readable text, intricate layering, or instructions the other models keep missing.
SDXL Inpainting comes later. If the outfit is approved but a collar, cuff, button line, or hem breaks, I mask that small region instead of rerunning the entire look. This keeps the approved identity, pose, and lighting intact and usually costs less than another full generation.
Run the fitting
For each garment, I generate three variations with the same prompt. The first pass is not a final image. It is a diagnostic.
I look at identity first, because that is the least negotiable part. Is it the same person, not merely a similar person? Are the face shape, hairline, skin tone, age, and expression stable? Are the hands anatomically believable? Then I look at fit. Does the garment hang from the actual shoulders and waist, or does it float like a sticker? Are the seams, buttons, straps, and closures in plausible positions?
Then I check fabric behavior. Leather should crease differently from cotton. Silk can catch light in a narrow band. Knitwear should follow the body without becoming plastic. A generated garment can look glossy and still be commercially useless if the material reads as vinyl.
When identity drifts, I do not stack more adjectives onto the prompt. I return to the source photo and choose a clearer front-facing image, repeat the identity sentence at the beginning, and reduce unrelated styling language. When the garment drifts, I improve the reference, add a back or detail image, and describe only the construction that actually distinguishes the piece. When one area is close but broken, I inpaint it.
Deliver the set
One image is rarely a fashion asset. I finish by making versions for the place where they will actually appear. A 4:5 crop works for social feeds, 3:4 is comfortable for catalogs and lookbook grids, and 9:16 gives me a still frame that can become short video later. I keep a filename convention such as look01-front, look01-detail, and look01-story, because a small campaign turns into a folder of mystery crops faster than I want to admit.
For a six-piece capsule, the credit math stays manageable. I can rehearse the prompt structure with the guest preview, then run three Seedream drafts per look at 5 credits each. If one approved image needs a collar repair, SDXL Inpainting adds 3 credits. I do not upgrade to a more expensive route until the fit and identity are already right.
Respect the person and the garment
This is the part I do not optimize away. I use a model photo only when the person has agreed to that use. I do not put a public figure's face into a lookbook. I do not generate children, school uniforms with sexualized framing, or suggestive content. I do not claim the image is a real try-on photo when it is a generated visualization.
Garments have owners too. If the piece belongs to another designer, I check whether I can reproduce and promote it. I also avoid asking a model to invent a protected logo. If branding is required, it belongs in a licensed design workflow, not in a improvised generation prompt.
For publishing, I keep product claims separate from creative imagery. A generated image can show styling direction. It should not invent fabric content, sizing claims, availability, or a testimonial from the person shown.
Common failures I now catch early
| Failure | What I change |
|---|---|
| Face becomes a similar stranger | Use a clearer identity reference, repeat identity first, and reduce scene instructions. |
| Body proportions shift | State that pose and proportions are locked, then reject any image that reshapes the body. |
| Garment silhouette changes | Add a cleaner front, back, or detail reference and describe the actual construction. |
| Hands melt into fabric | Choose a source pose with hands visible and away from critical seams. |
| Closures multiply or vanish | Generate a detail crop, then repair the smallest possible mask. |
| Lighting looks pasted on | Keep the original light direction and name its direction in the prompt. |
| Skin becomes porcelain | Ask for natural skin texture and reject over-retouched results. |
Frequently asked questions
Can I use one model photo for an entire collection?
Yes, if you have permission and keep the identity reference stable. I reuse the same model photo and change only the garment reference, so looks remain comparable.
How many garment references do I need?
One clean front view can be enough for a simple T-shirt. Tailored clothing, dresses, and pieces with hardware benefit from a back view or detail shot.
Which model should I start with?
For commercial photoreal delivery, start with Seedream 5.0 Lite and multiple references. Move to GPT Image 2 for complex layering or readable details, and use SDXL Inpainting for local repair.
Can I publish AI outfit swaps as real photos?
Only if the destination explicitly allows it and the claim is truthful. Most commercial work should label generated imagery when required and avoid implying a physical fitting occurred.
How do I keep the cost predictable?
Rehearse with the guest preview, generate three drafts per garment, and repair one approved image locally instead of rerunning the whole look.