AI Photo Editor for Ecommerce: What It Costs and Which Model to Use
Most guides tell you to “use AI for product photos” and stop there. This one gives you the actual per-image cost on each model, the resolution multiplier that quietly doubles your bill, and which shot to generate first.
July 23, 2026 · 6 min read
Contents
Quick answer
For ecommerce product photos, draft at standard resolution on a cheap model, then re-run only your keeper at high resolution. On Shotari that means drafting with Seedream 5 Lite at 3 credits or GPT Image 2 at 5 credits, then finishing the winner at 4K — where the same generation costs twice as much. Drafting cheap and finishing high is the whole trick, and it typically cuts the credits per usable image by more than half.
The one number to remember
Resolution multiplies the model's base cost: standard is ×1.0, 2K is ×1.5, and 4K is ×2.0. Portrait and landscape aspect ratios add another ×1.1 on top.
What actually breaks in product photos
Product photography fails in a small number of specific ways, and knowing which one you are fighting decides which tool you reach for. Blanket “enhance my photo” prompts waste generations because they ask the model to fix everything at once.
The background is the problem
Cluttered desk, wrong colour, visible edge of the seamless paper. This is a background replacement job, not a generation job — start from your real photo as a reference so the product itself is untouched.
The lighting is the problem
Flat phone-flash lighting, harsh shadow under the product, blown highlights on glossy surfaces. Describe the light setup explicitly: softbox direction, fill, and whether you want a contact shadow.
The context is the problem
The product is fine but sits in a void. Lifestyle framing — on a marble counter, in a hand, on a linen surface — is what marketplace listings and ads actually convert on.
The resolution is the problem
The shot is right but too small for a zoom view. This is the only case where paying the 4K multiplier upfront makes sense.
Which model, and what it costs
These are the real base costs per generation at standard resolution on Shotari. There is no editorial ranking here — the cheap models are genuinely fine for drafting, and the expensive ones earn their price only on the final pass.
| Model | Credits (standard) | Best used for |
|---|---|---|
| Shotari Basic | 0 | Free drafting. Slow queue, text-to-image only — good for exploring composition before you spend anything. |
| Shotari Pro | 2 | Cheap iteration when Basic is not sharp enough. |
| Seedream 5 Lite | 3 | The cheapest premium model. Strong default for high-volume product drafts. |
| Nano Banana | 4 | Reliable background and scene replacement from a reference photo. |
| GPT Image 2 | 5 | Anything with text in frame — packaging, labels, on-image copy. |
| GPT Image 1.5 | 5 | Faster OpenAI sibling; clean text, useful when GPT Image 2 is busy. |
| Seedream 4.5 | 6 | Detailed product surfaces and fabric texture. |
| Nano Banana 2 | 6 | General-purpose upgrade over Nano Banana. |
| FLUX.2 | 6 | Stylised and editorial-looking product shots. |
| Nano Banana Pro | 10 | The finishing pass when the shot is going on your product detail page. |
Free credits — the 20 you get at signup and the 10 from each daily check-in — only run Shotari Basic and Shotari Pro. Every model above those two needs paid credits, which is why a non-zero balance can still report “insufficient credits” the moment you switch models.
The resolution multiplier trap
The number on the model card is the standard-resolution price. The actual charge is that base cost multiplied by your resolution and aspect ratio, then rounded up.
| Setting | Multiplier |
|---|---|
| Standard (1K) | ×1.0 |
| 2K | ×1.5 |
| 4K | ×2.0 |
| Square, classic, tall ratios | ×1.0 |
| Portrait (9:16) or landscape (16:9) | ×1.1 |
So a GPT Image 2 generation at 4K in 16:9 is not 5 credits — it is 5 × 2.0 × 1.1, rounded up to 11. A Nano Banana Pro shot at 4K square is 20. Run twelve of those while you are still deciding on the composition and you have spent 240 credits to answer a question that Shotari Basic answers for nothing.
Practical rule
Never change resolution and prompt in the same step. Lock the composition at standard resolution first, change nothing but the resolution on the final run.
The four-shot workflow
Marketplace listings and ad sets need the same four shots almost every time. Generating them in this order means each one reuses what the last one established.
1. The clean pack shot
Product on a plain seamless background, centred, with a soft contact shadow. Upload your real photo as a reference so the product geometry and colour stay true. This shot is your listing thumbnail and the reference for everything after it.
2. The detail shot
Close crop on the material, stitching, texture, or label. This is where Seedream 4.5 earns its 6 credits and where GPT Image 2 is worth it if there is readable text on the packaging.
3. The lifestyle shot
Product in a real setting — kitchen counter, desk, held in a hand. This is the shot that carries ads and social. Describe the surface and light, not just the scene.
4. The scale shot
Product next to a common object so buyers understand its size. The single most underused shot in ecommerce, and it prevents a large share of size-related returns.
You can attach up to three reference images to a single generation. Use one for the product, one for the lighting or mood you want, and one for the background style — that combination is far more controllable than trying to describe all three in words.
Prompts that work
Product prompts fail when they describe the product instead of the photograph. The model can already see the product in your reference image; what it needs from you is the camera, the light, and the surface.
Clean pack shot
The product from the reference image on a seamless off-white background, centred, soft large softbox from the upper left, gentle contact shadow beneath, no props, shot on 85mm at f/8, even studio exposure
Detail / texture
Macro close-up of the product surface from the reference image, raking side light to reveal texture, shallow depth of field, neutral colour, no background distraction
Lifestyle
The product from the reference image on a warm oak kitchen counter, morning window light from the right, soft natural shadows, a linen cloth slightly out of focus behind, lifestyle editorial look
Scale reference
The product from the reference image standing next to a standard coffee mug on a plain light grey surface, straight-on eye-level camera, even diffuse lighting, both objects fully in frame and in focus
Mistakes that cost you credits
- ⚠
Drafting at 4K
Doubles every exploratory generation. Explore at standard resolution, upgrade once.
- ⚠
Changing two things at once
If you edit the prompt and the model in the same step, a worse result tells you nothing about which change caused it. Change one variable per run.
- ⚠
Describing the product instead of the photo
With a reference image attached, words spent re-describing the product are wasted. Spend them on light, surface, lens, and framing.
- ⚠
Naming real brands in the prompt
Prompts naming real brands or logos get screened out before the model runs. Describe the design instead — the request bounces back immediately otherwise.
- ⚠
Regenerating instead of referencing
If shot two should match shot one, feed shot one back in as a reference image rather than hoping the same prompt reproduces it. Generations are not deterministic.
Try it on your own product photo
Upload a product shot, start on the free tier to lock the composition, then finish the keeper on a premium model at high resolution.
Open the product photo editor