Why AI Photo Edits Fail โ 3,231 Production Runs
Text-to-image succeeded 98.7% of the time. Image-to-image, on the same model and the same infrastructure, succeeded 83.6%. The gap is not speed and it is not the prompt.
September 11, 2026 ยท 6 min read
On this page
Quick answer
Across 3,231 generations from 572 distinct users between 14 July and 11 September 2026, text-to-image succeeded 98.7% of the time and image-to-image succeeded 83.6% โ a 15-point gap on the same model, same pipeline, same median latency. The edits do not fail because editing is slower or because the instructions are worse. They fail because a second input exists: the uploaded photo. Every single content refusal in the dataset โ 169 of them โ happened on the image-to-image path. Text-to-image produced zero. If your edit is being refused, the thing to change first is the photo, not the wording.
Why publish this
Public discussion of AI edit failures is almost entirely prompt advice, because prompts are the part users can see. We run the pipeline, so we can see the other side: which requests reached the model, what came back, and how long it took. Numbers below are read straight from our generation table, not estimated. Sample window and limits are stated at the end.
The measurement
| Path | Runs / success rate | Median time |
|---|---|---|
| Text-to-image | 466 ยท 98.7% | 82s |
| Image-to-image | 2,765 ยท 83.6% | 81s |
| All runs | 3,231 ยท 85.8% | 81s |
The latency line is the one that surprises people. Editing an existing photo is widely assumed to be the slower operation โ it is not. Median time to a finished image is within one second of text-to-image. What differs is the failure rate, and failures are not evenly shaped either.
| Failure mode | Image-to-image | Text-to-image |
|---|---|---|
| Content refusal | 169 (37.3%) | 0 |
| Timed out before an image | 150 (33.1%) | 3 |
| Model replied with text, no image | 131 (28.9%) | 2 |
| Other | 3 (0.7%) | 1 |
Refusals come from the photo, not the prompt
169 refusals on the editing path. Zero on the text path. That asymmetry is the most useful thing in this dataset, because it contradicts how the problem is usually discussed. If refusals were driven by instruction wording, they would appear on both paths โ the same users write the same kind of instructions either way. They appear on exactly one.
Reading the refusals, three categories cover nearly all of them: photographs of real identifiable people where the requested change touches the body or the face; recognisable copyrighted characters; and anything shaped like an identity document, card or official paper. In each case the moderation decision is made once the image is in context โ which is why the same sentence that works on one photo fails instantly on another.
The fastest diagnostic
Run the same instruction on a different photo. If it succeeds, the instruction was never the problem, and rewriting it is wasted effort. This one test separates a refusal from a genuine prompt problem in about 90 seconds.
When the model answers in words instead of pixels
131 runs came back as prose: a paragraph describing what it would do, or explaining why it would rather not, with no image attached. This is not an outage and not a refusal โ the request completed, it simply produced the wrong kind of output. It clusters around instructions that read as questions or as open-ended requests ("can you make this look better?"), where a language model's most natural response really is a sentence.
Instructions phrased as a concrete visual edit โ naming what changes and what must not โ do not produce this failure at anywhere near the same rate. That is the one place where wording genuinely matters, and it matters for a different reason than people assume: not to persuade the model, but to make the output type unambiguous.
Timeouts are a budget problem, not a speed problem
150 image-to-image runs never returned an image inside the window. We spent real effort attributing these to the model being slow before measuring properly: with median completion at 81 seconds and p90 at 111 seconds, most work finishes well inside two minutes. The tail is long, not the middle. A run sitting at 30 seconds is not stuck โ and a timeout set from average behaviour will cut off the runs that were about to succeed.
We learned this the expensive way: an earlier version of our pipeline timed out downloads at 20 seconds while the full-size file genuinely took around 28 to prepare. Every one of those was a completed generation thrown away at the last step. End-to-end timing cannot locate a single-stage timeout โ you have to instrument each stage separately.
What actually reduces failures
1. Change the photo before you change the prompt
Refusals are input-driven. If an edit is declined, the highest-value next action is a different source image โ a different crop, a different subject, or the same scene without a face or document in frame. Rewriting the sentence first is the common move and the least effective one.
2. Name the change and the constraint in the same instruction
"Change only the background; keep the person's face, hairstyle and proportions exactly as they are." Naming what must not change does two jobs: it reduces drift in the parts you did not mention, and it makes the expected output unmistakably an image rather than an answer.
3. One change per run
Every run is a fresh render rather than a layer applied to the last one, so two instructions in a single prompt compound drift instead of adding up. Make the change, look at it, run again from the result.
4. Give it more than 30 seconds
Median 81 seconds, p90 111. Judge a run at two minutes, not at thirty seconds. Retrying early wastes an attempt that was likely to land.
If you want to test the four points above against your own photos, the GPT Image 2 editor runs the same pipeline these numbers came from. The built-in Shotari Basic model costs 0 credits, so the diagnostic in the callout above โ same instruction, different photo โ is free to run. Related reading: GPT Image 2 not working for the error-by-error breakdown, and how to write AI image prompts for the instruction side.
Method and limits
- โ
What the numbers are
Every generation recorded by our pipeline between 2026-07-14 and 2026-09-11: 3,231 runs, 572 distinct users, counted from the generation table with no sampling and no exclusions. Success means a finished image was produced and delivered. Durations are server-side, from job start to finish.
- โ
What they are not
This is one pipeline and predominantly one model family, so the absolute rates are ours, not the industry's. The shape of the finding โ refusals concentrated entirely on the image path, latency effectively identical between paths โ is what we would expect to generalise; the exact percentages are not.
- โ
Where the sample is thin
Text-to-image is 466 runs against 2,765 for editing, because our traffic is overwhelmingly people editing a photo they already have. The 98.7% figure therefore carries a wider error bar than the 83.6% one. We are not reporting per-model comparisons at all: our second model has only 23 runs in this window, which is far too few to publish a rate for.
- โ
Reuse
These figures may be quoted with attribution to Shotari. If you want a cut we did not publish โ by month, by failure string, by resolution โ ask and we will run it.
Test it on your own photo
Same instruction, different photo โ the diagnostic that separates a refusal from a prompt problem. Free on Shotari Basic, no account needed for the first runs.
Open the editor