AI Model Outfit Swap Tutorial: How I Keep the Person and Change the Clothes

A first-hand workflow for AI model outfit swaps, from consented source photos and garment references to model routing, failure checks, and commercial delivery.

Published: 2026-09-01 - By imgmov Team

The first outfit swap I liked was still wrong. The jacket was convincing, the light matched, and then I looked at the collar. The model's face was fine, but the garment had become a vague cousin of the original piece. That test taught me the most useful rule for an AI model outfit swap. The person is an identity problem, but the garment is a product-accuracy problem. They need different evidence.

Since then, I run every swap as a small fitting session inside imgmov. I upload a consented model photo, add clean garment references, keep the pose and background locked, and only then let the model change clothing. The workflow below is the version I now reuse when I need a set of lookbook-style images without reshooting every piece.

Prepare the evidence first

A good source model photo answers most of the workflow before generation starts. I look for a full-body adult subject, plain studio light, a pose I can repeat, and clothing that does not hide the silhouette I need to replace. Hands matter more than I expected. If the fingers disappear into a pocket or cross the garment seam, the model has to invent fabric around them.

My pre-flight list is short:

The clothing reference is just as important. A front view is rarely enough for a tailored jacket, wrap dress, pleated skirt, or anything with hardware. I prefer a front reference plus a back or detail shot when the garment has closures, embroidery, a distinctive neckline, or a special hem. I crop away mannequins, hangers, floors, and price tags when possible. The model should receive the garment, not a room full of retail furniture.

Define the two roles

My prompt starts by naming the references. Otherwise the model may average the two images into a new person wearing an imagined outfit.

Reference 1 is identity. Keep the same adult model, face, hairstyle, skin tone, body proportions, pose, camera angle, framing, and background.
Reference 2 is garment only. Replace the current outfit with the exact garment in Reference 2: [color, material, neckline, sleeves, closure, length, fit].
Preserve natural fabric folds, seams, shadow direction, and skin texture. Do not beautify the face, reshape the body, add a different person, add text, or add a watermark.

That phrase, "garment only," does a lot of work. It separates the face I want to preserve from the piece I want to move. If I am swapping a top and keeping trousers, I also state the trousers explicitly. A model will otherwise decide on its own.

For a first test, I keep everything constant except one variable. The same model photo, same camera framing, same background, same prompt structure. I change only the garment reference. If I change pose, scene, fabric, and accessory at once, I cannot tell which instruction caused the failure.

Choose the model by failure mode

On imgmov, I no longer ask which model is best in general. I ask which failure I can afford to fix.

ModelWhen I use itWhy
Agnes Image 2.1 FlashGuest rehearsal onlyIt supports image-to-image, but guest previews are watermarked, limited to three per day, and not for commercial delivery.
Seedream 5.0 LitePrimary photoreal outfit swapSupports multiple image references and costs 5 credits per image.
GPT Image 2Complex layering, readable labels, or graphic garment detailsSupports multiple image references and costs 2-80 credits by quality and resolution.
SDXL InpaintingOne damaged area in an otherwise approved imageCosts 3 credits per image and works with a mask.
Nano-Banana 2Text-only concept boards, not the exact swapIt is useful for exploring wardrobe mood before I spend credits on production references.

Agnes is my rehearsal layer when I am logged in as a guest. It can tell me whether the prompt roles are clear, but I never ship a watermarked preview. For the production route, Seedream 5.0 Lite is usually my first choice because it handles multiple references and keeps photographic context stable. GPT Image 2 earns its credits when a look includes readable text, intricate layering, or instructions the other models keep missing.

SDXL Inpainting comes later. If the outfit is approved but a collar, cuff, button line, or hem breaks, I mask that small region instead of rerunning the entire look. This keeps the approved identity, pose, and lighting intact and usually costs less than another full generation.

Run the fitting

For each garment, I generate three variations with the same prompt. The first pass is not a final image. It is a diagnostic.

I look at identity first, because that is the least negotiable part. Is it the same person, not merely a similar person? Are the face shape, hairline, skin tone, age, and expression stable? Are the hands anatomically believable? Then I look at fit. Does the garment hang from the actual shoulders and waist, or does it float like a sticker? Are the seams, buttons, straps, and closures in plausible positions?

Then I check fabric behavior. Leather should crease differently from cotton. Silk can catch light in a narrow band. Knitwear should follow the body without becoming plastic. A generated garment can look glossy and still be commercially useless if the material reads as vinyl.

When identity drifts, I do not stack more adjectives onto the prompt. I return to the source photo and choose a clearer front-facing image, repeat the identity sentence at the beginning, and reduce unrelated styling language. When the garment drifts, I improve the reference, add a back or detail image, and describe only the construction that actually distinguishes the piece. When one area is close but broken, I inpaint it.

Deliver the set

One image is rarely a fashion asset. I finish by making versions for the place where they will actually appear. A 4:5 crop works for social feeds, 3:4 is comfortable for catalogs and lookbook grids, and 9:16 gives me a still frame that can become short video later. I keep a filename convention such as look01-front, look01-detail, and look01-story, because a small campaign turns into a folder of mystery crops faster than I want to admit.

For a six-piece capsule, the credit math stays manageable. I can rehearse the prompt structure with the guest preview, then run three Seedream drafts per look at 5 credits each. If one approved image needs a collar repair, SDXL Inpainting adds 3 credits. I do not upgrade to a more expensive route until the fit and identity are already right.

Respect the person and the garment

This is the part I do not optimize away. I use a model photo only when the person has agreed to that use. I do not put a public figure's face into a lookbook. I do not generate children, school uniforms with sexualized framing, or suggestive content. I do not claim the image is a real try-on photo when it is a generated visualization.

Garments have owners too. If the piece belongs to another designer, I check whether I can reproduce and promote it. I also avoid asking a model to invent a protected logo. If branding is required, it belongs in a licensed design workflow, not in a improvised generation prompt.

For publishing, I keep product claims separate from creative imagery. A generated image can show styling direction. It should not invent fabric content, sizing claims, availability, or a testimonial from the person shown.

Common failures I now catch early

FailureWhat I change
Face becomes a similar strangerUse a clearer identity reference, repeat identity first, and reduce scene instructions.
Body proportions shiftState that pose and proportions are locked, then reject any image that reshapes the body.
Garment silhouette changesAdd a cleaner front, back, or detail reference and describe the actual construction.
Hands melt into fabricChoose a source pose with hands visible and away from critical seams.
Closures multiply or vanishGenerate a detail crop, then repair the smallest possible mask.
Lighting looks pasted onKeep the original light direction and name its direction in the prompt.
Skin becomes porcelainAsk for natural skin texture and reject over-retouched results.

Frequently asked questions

Can I use one model photo for an entire collection?

Yes, if you have permission and keep the identity reference stable. I reuse the same model photo and change only the garment reference, so looks remain comparable.

How many garment references do I need?

One clean front view can be enough for a simple T-shirt. Tailored clothing, dresses, and pieces with hardware benefit from a back view or detail shot.

Which model should I start with?

For commercial photoreal delivery, start with Seedream 5.0 Lite and multiple references. Move to GPT Image 2 for complex layering or readable details, and use SDXL Inpainting for local repair.

Can I publish AI outfit swaps as real photos?

Only if the destination explicitly allows it and the claim is truthful. Most commercial work should label generated imagery when required and avoid implying a physical fitting occurred.

How do I keep the cost predictable?

Rehearse with the guest preview, generate three drafts per garment, and repair one approved image locally instead of rerunning the whole look.

Related reading

Back to blog | Pricing | Start creating