GPT Image 2 is a next-generation AI image model: more frame formats, resolution up to 4K and genuinely useful reference handling. Upload as many as 16 of your own photos, describe the job in plain words, and the model edits a shot or builds a new scene from your sources. Everything runs in the browser with a full generation history.

The upgrade widens the canvas literally: six aspect ratios instead of three, including vertical 9:16 for stories and wide 16:9 for covers. Resolution is now a choice — from quick 1K to dense 4K. And an Auto mode picks the proportions for you, which is a quiet lifesaver when editing your own photos.
The line's signature trait survives: careful prompt reading. The model still holds long instructions with many conditions, but executes them with finer detail. If you used the previous generation, the honest test is rerunning an old prompt and comparing how the small stuff is drawn.
Attach up to 16 images of 10 MB each. One photo plus an instruction means editing: remove passers-by, replace the sky, repaint a facade. Several photos mean assembly: the product from one shot, the interior from another and the mood from a third merge into a new frame under your direction.
A practical retail workflow is one photoshoot for the whole catalog: shoot the product once, then generate it in different settings, on models, in seasonal scenes. For personal photos it rescues a good shot with a bad background. Describe the change in words, like talking to a retoucher — no masks or layers.

You get square 1:1, vertical 9:16 and 3:4, and horizontal 16:9 and 4:3 — enough to cover stories, feeds, video thumbnails, banners and print layouts. Choose the format before generating; picking the right frame up front beats cropping a finished image later.
Auto mode removes the decision entirely: the model judges which proportions fit. With a reference attached this is especially handy — the framing follows the source, so a vertical photo never gets squeezed into a square. When a platform demands an exact size, just set it explicitly.
1K is the workhorse for drafts and idea checks: it generates faster and spends fewer tokens. 2K is the all-rounder for social feeds and websites viewed on screens. 4K earns its cost on print jobs, large banners and images people will zoom into or examine up close.
The economical routine is two-step: hunt for the right composition and wording at 1K, then regenerate the winner at 4K using the same prompt from your history. You avoid paying for high resolution during experiments and keep full quality where it actually matters.
Marketplace sellers get consistent product cards without a studio. Social media managers get covers and stories cut exactly to platform size. Designers get quick concepts for interiors, packaging and posters where both brief accuracy and output resolution count.
Beginners are forgiven a lot here: there is no special prompt language to learn, ordinary sentences work. Start with a simple task — swapping the background on your own photo is a good first run — check the result in history, and work up to multi-reference scene assembly.
The newer model offers six frame formats instead of three, including 9:16 and 16:9, adds an Auto aspect mode and three resolution tiers up to 4K. The trademark comprehension of complex prompts remains, with noticeably finer detail.
Yes. Upload the photo as a reference and state the change in words: remove an object, replace the background, recolor the outfit. No mask is needed — text instructions are enough, and you can combine up to 16 images in one run.
The model picks frame proportions to match your prompt or your uploaded reference. When editing photos this keeps the original's natural framing intact. If a platform requires an exact size, select the specific ratio manually instead.