OpenAI won’t say what’s inside GPT-Image-2.5. They don’t have to. The public API still gives us plenty to work with: pricing, resolution constraints, how output is metered, and how references behave.
Add in the research history of Boyuan Chen, one of the researchers training GPT image generation, and I think there’s a stronger possibility: The question no longer has a clean, binary answer.
Images 2.5 has some obvious improvements that everyone noticed: It preserves reference subjects better, confines edits more reliably to what you asked it to change, and holds onto earlier changes across multiple turns. According to OpenAI, generation latency is also down by as much as 50% compared with Images 2.0.
What OpenAI still won’t discuss is the architecture. When Images 2.0 launched in April, research lead Boyuan Chen described it as “revamped from scratch,” and declined to say whether it was diffusion or autoregressive. The Images 2.5 announcement doesn’t resolve the question either; OpenAI talks about capabilities instead.
So I want to approach it from another direction.
Not “What did OpenAI build?”
But “What would have to be true for the API to behave like this?”
Four clues hiding in the API
Sunburst and Flare look like model variants, not a pipeline
OpenAI calls Flare its small model and Sunburst its base model. Flare is optimized for speed; Sunburst is optimized for quality.
That doesn’t tell us how the two models are related internally, but “small” and “base” sound much more like a capacity relationship than a pipeline relationship.
If I had to bet, Flare is distilled from Sunburst, compressed from the same underlying family, or related in some comparably direct way. I would not bet that Sunburst is an agentic wrapper that calls Flare repeatedly and gets quality from pipeline depth.
Vector-quality output is still raster
The API outputs PNG, JPEG, or WebP. Transparency is available through PNG and WebP. SVG is not part of the response format.
That does not prove there are no vector-like or structured representations inside the model. It only tells us that whatever happens internally has been rasterized by the time it reaches us.
Still, the result is striking. Some generated logos have edges clean enough to make you wonder whether there’s a Bézier curve hiding somewhere upstream. Zoom in far enough, though, and you get antialiasing and slightly imperfect curve continuity, rather than mathematical primitives.
Multi-reference looks native, not bolted on
The editing API accepts up to 16 reference images, and can optionally take a mask for localized editing.
My guess is that reference images enter the same general representational world the model uses for generation, rather than being handed off to a bespoke reference-matching subsystem. Maybe they become image tokens. Maybe they land in a closely aligned latent space.
If that’s right, the reason multi-reference works is that it isn’t special. If it’s native context rather than a bolt-on feature, the model isn’t “consulting” 16 external images; it’s generating with 16 images already in the problem.
Token billing is a thread worth pulling
Both 2.5 models cost $30 per million output image tokens, and both expose low, medium, high, xhigh, and max quality settings. Models and quality settings can consume different numbers of output tokens.
A conventional fixed-step latent diffusion sampler doesn’t naturally present itself as a variable-length output sequence. Its compute is easier to think about in terms of latent area and denoising work.
That doesn’t necessarily mean “token billing = autoregressive model.” OpenAI could simply be pricing image compute in token-shaped units so it fits cleanly into the rest of the API. But a look at older token tables makes that harder to believe.
Useful information in the old token table
For GPT Image models before GPT-Image-2, OpenAI published the number of specialized image tokens required at different resolutions and quality levels. The docs are explicit that these numbers belong to older models and this is not a token map for Sunburst or Flare.
But look at the arithmetic.
For those models, a 1024 × 1024 image used 272 output tokens at low quality, 1,056 at medium, and 4,160 at high. A 1024 × 1536 image used 408, 1,584, and 6,240. The numbers factor strangely neatly:
| Resolution | Quality | Output tokens | Factorization | Grid |
|---|---|---|---|---|
| 1024 × 1024 | low | 272 | 16 × 17 | 16 rows, 16 columns |
| 1024 × 1024 | medium | 1,056 | 32 × 33 | 32 rows, 32 columns |
| 1024 × 1024 | high | 4,160 | 64 × 65 | 64 rows, 64 columns |
| 1024 × 1536 | low | 408 | 24 × 17 | 24 rows, 16 columns |
| 1024 × 1536 | medium | 1,584 | 48 × 33 | 48 rows, 32 columns |
| 1024 × 1536 | high | 6,240 | 96 × 65 | 96 rows, 64 columns |
Every entry is rows × (columns + 1).
The overhead grows with the number of rows, not with total image area. One natural explanation is a raster-style sequence with a boundary marker at the end of each row: Emit the spatial positions, emit a delimiter, wrap.
I can’t see the tokenizer, so “an end-of-row” token is just a hypothesis. It could also be codec bookkeeping, or some artifact of how OpenAI counted the older representation. But awkward constants can be more useful than clean abstractions when you’re trying to infer an implementation. A global header should look global. A cost that increments once per row looks like something happening once per row.
There’s a second coincidence hiding in the same table. At high quality, 1,024 pixels maps to 64 spatial positions. That’s 16 pixels per position.
GPT-Image-2.5, meanwhile, requires both output dimensions to be divisible by 16.
This doesn’t prove that 2.5 inherited the old image tokenizer. OpenAI said they rebuilt the architecture, so it would be surprising if everything stayed the same. But the old model exposes a 16-pixel spatial granularity, and the new model exposes a 16-pixel divisibility constraint. That’s consistent with an inherited tokenizer, but it’s also the standard constraint for latent diffusion transformers, so on its own it doesn’t discriminate.
The minimum size looks like a model limit
The rest of the 2.5 size contract is unusually specific: Both edges must be multiples of 16, the aspect ratio can’t exceed 3:1, neither edge can exceed 3,840 pixels, and total area has to fall between 655,360 and 8,294,400 pixels.
If you divide the two pixel-count bounds by 256 (the area of a 16×16 patch) you get 2,560 and 32,400 possible patch locations.
The lower number is suspiciously tidy. But there’s an obvious caveat: When both dimensions have to be divisible by 16, every legal canvas will have an area divisible by 256. So the fact that the bounds divide cleanly is not, by itself, evidence of a 16 × 16 tokenizer.
The minimum area rule is stranger. A 640 × 1024 image hits the minimum exactly and is legal. A 512 × 512 image is not, even though 512 is a perfectly ordinary image dimension. Nothing is wrong with either edge of the square, but the whole canvas contains too few pixels.
A minimum useful sequence length would produce that sort of rule. So would a minimum latent canvas, a hole in the training distribution, or a product decision to suppress low-quality outputs.
The 3:1 cutoff points to a whole-canvas model
Suppose the core generator were primarily agentic: Plan a canvas, generate independent pieces, stitch or composite them together. An extreme banner shouldn’t be fundamentally strange under that architecture. A 10:1 web banner is easy to tile: generate a few panels, blend the seams, done.
Instead, the API stops at 3:1.
That makes more intuitive sense if most or all of the canvas passes through one model as a coherent spatial object. Keep stretching the grid and eventually positional representations, attention geometry, or simply the training distribution stop behaving well.
This isn’t proof against tiling or layered generation; OpenAI could have perfectly mundane product reasons for banning extreme ratios. But if you ask which underlying architecture makes the constraint feel least arbitrary, I would choose the whole-canvas model.
The support for arbitrary legal dimensions points in the same direction. Rigid absolute position embeddings would be awkward here unless they were being interpolated. A factorized or 2D rotary positional scheme, trained across multiple resolutions and aspect-ratio buckets, would fit the surface behavior much more naturally.
What xhigh and max can tell us about the representation
If some descendant of the old spatial representation survives in 2.5, xhigh and max have two obvious places to spend additional token budget.
One is more locations. In the old table, each step up in quality made the grid finer: 64 pixels per position at low, 32 at medium, 16 at high. Continue the pattern and xhigh becomes an 8-pixel grid.
The other is more information per location: residual codes at the same grid points, additional quantization planes, another level in a coarse-to-fine hierarchy, or some other refinement representation.
Either would fit what higher quality is for. Once composition is established, the remaining wins are things like small type, fine edges, material detail, and texture. I lean toward the second, mostly because of scale: an 8-pixel grid at 3840 × 2160 would mean about 130,000 positions for a model that, on my reading, handles the whole canvas at once.
We no longer have to settle this from old pricing tables; OpenAI exposes token usage for 2.5, and recommends measuring it because consumption varies by model, size, and quality. The two options predict differently shaped counts (see the experiments below). If neither shape appears, the old structure probably didn’t survive.
Diffusion vs. autoregression may be the wrong axis
Here’s the thing that makes the original question unanswerable rather than just unanswered: The two paradigms were never orthogonal, and the field has spent years collapsing them into one another.
Take a sequence of tokens and a denoising timeline. In ordinary full-sequence diffusion, the whole sequence moves through broadly the same noise schedule together. In an autoregressive model, the past is already resolved while generation advances into the future.
Those sound like fundamentally different processes until you allow each token to have its own noise level.
That’s the idea behind Boyuan Chen’s research paper “Diffusion Forcing.” The model is trained to denoise sequence tokens with independent per-token noise levels, so one part of the sequence can be clean while another is still heavily corrupted. The same training framework can then support behavior that looks more like full-sequence diffusion, behavior that looks more like next-token generation, and schedules between the two.
There’s no public evidence that GPT-Image-2.5 uses Diffusion Forcing.
There doesn’t need to be for Chen’s research to matter here. It establishes something narrower and more important: Modern generative systems truly blur the lines between diffusion and autoregression.
Chen’s research explored this boundary
Chen is a research scientist at OpenAI who describes himself as “one of the few researchers training GPT image generation.” His selected work includes Diffusion Forcing, SpatialVLM and History Guidance.
On his “Diffusion Forcing” project page, Chen suggests a conditioning setup where context tokens stay clean while future tokens remain noisy. Translate that idea from a sequence into an image: “Keep this region, change that one.” Sound familiar?
If some representation of the region you want preserved can stay clean while another region is allowed to move, localized editing no longer requires the model to destroy and reconstruct the entire image on every turn. I’m not claiming that’s how GPT-Image works, and keeping context clean isn’t unique to Diffusion Forcing: ordinary diffusion inpainting already holds known regions fixed while regenerating the rest. What Diffusion Forcing adds is training where mixed clean and noisy states are the normal case rather than a special editing mode. That’s a natural fit for multi-turn editing, and it comes from a researcher who now trains GPT image models.
Chen’s follow-up work, “History-Guided Video Diffusion,” targets a related problem from the temporal side: conditioning on different amounts of history while improving long-horizon consistency.
Obviously, a long video rollout is different from a 12-turn image editing chain. Still, the structural problem rhymes: How much previous state can you keep intact while continuing to generate?
The boundary was already disappearing
Chen’s paper gave the idea a clean frame, but diffusion and autoregression had been bleeding into one another for years before he published his research:
In 2021, D3PM connected discrete diffusion with autoregressive and mask-based generation, while Hoogeboom et al.’s “Autoregressive Diffusion Models“ put order-agnostic autoregression and absorbing-state diffusion inside a broader model class and TimeGrad used a diffusion model inside an autoregressive forecasting system.
In 2023, “AR-Diffusion” gave different token positions different denoising schedules
In 2024, “Rolling Diffusion” varied noise progressively across a temporal sequence.
Then, hybrid designs became harder to classify at all. MAR uses diffusion to model continuous per-token distributions inside an autoregressive image generator. Transfusion trains one transformer with next-token prediction for text and diffusion for images. DART explicitly unifies autoregressive modeling and diffusion while denoising image patches. CausalFusion factorizes generation across both sequential tokens and diffusion noise levels. Show-o combines autoregressive text modeling with discrete diffusion for images, while JanusFlow pairs an autoregressive language model with rectified flow.
At that point, asking whether an image model “is diffusion” starts to sound like asking if a hybrid car is gas or electric. The answer is yes, but…
So why are the edits so precise?
I have three guesses:
References probably live close to the representation being generated
If reference images are encoded into the same vocabulary, or into a latent space closely aligned with the model’s output representation, then “preserve this” is a very different task from resynthesizing it from a text description.
You have source structure available directly.
That would fit both the multi-reference API and the way OpenAI describes Images 2.5: Reference subjects are more recognizable, targeted edits disturb less of the surrounding image, and earlier work survives more reliably across subsequent edits.
I don’t think each edit starts from only the last decoded image
To make edit chains degrade, decode an image, re-encode that output on the next turn, decode again, and repeat. Every round is another opportunity to lose detail.
OpenAI’s Responses API supports keeping image generation outputs and image IDs in the conversation context, or continuing through previous_response_id.
That doesn’t reveal what the image model itself conditions on, but paired with the multi-turn consistency claim, it hints at conditioning on the original plus the edit history.
You can grade text rendering
People tend to overthink architectural explanations when image models suddenly get much better at text. I would start with the boring reason: data.
Synthetic supervision for text rendering is cheap and exact. Render arbitrary strings in arbitrary fonts at known locations and you know the answer before the model sees the example. Glyph identity is labeled. Position is labeled. Bounding boxes are available.
Then add OCR and you have something close to an automatic verifier for whether the model spelled the word correctly.
This still isn’t solved; OpenAI’s own docs list precise text placement and clarity among GPT Image’s limitations. But the point here is that text has a reward signal that “make this beautiful” doesn’t, which suggests a broader pattern that’s more important than the bigger questions about model architecture: The easiest capabilities to push hard are often the ones you can check automatically.
If you can build a cheap verifier, you can put it inside a training loop, an inference loop, or both. That idea shows up over and over in the systems we’re building, which is why it matters more than the question of whether GPT-Image-2.5 uses one type of architecture or another.
Experiments that could prove this wrong
Most of these guesses are cheap to test experimentally:
Diff the untouched region. Make a tightly scoped edit, then compare the rest of the output with the source. If untouched regions are bit-identical under lossless output, there’s probably an explicit copy or composite operation. If they’re extremely close but not identical, it would point toward regeneration under strong conditioning.
Fit the 2.5 token formula. Request a spread of legal dimensions at every quality tier and inspect usage. Fit the old rows × (cols + k) pattern, pure area, piecewise functions, or whatever else survives contact with the data. If a per-row term appears again, the raster ordering hypothesis gets stronger. If it doesn’t, it’s less likely the old raster-ordering structure survived into the new model.
Regress latency against tokens and pixels separately. If wall-clock time follows output token count better than total image area, that supports sequence-heavy decoding. A pattern closer to repeated whole-canvas refinement points elsewhere.
Probe high → xhigh → max. Measure all three tiers at several resolutions and try to factor the counts the way the old table factors. If xhigh comes out as rows × (cols + 1) with twice high’s rows and columns (128 × 129 = 16,512 at 1024 × 1024, if high is still 64 × 65), the grid got finer. If it factors around high’s grid instead, with the same rows carrying more tokens each or high’s whole grid repeated, the extra budget went into each location. Counts that factor neither way would suggest something else.
Where I’d actually put my money
My guess is:
A large multimodal transformer emitting a coarse-to-fine or masked-parallel token sequence, with a continuous (probably diffusion-style) head producing the values at each position, trained at native resolution with 2D rotary positions, and heavily distilled for the small variant.
Why this combination?
Because the capability profile itself looks hybrid.
Spelling, logos, geometry, identity, and composition benefit from global structure and internal consistency. Hair, lighting, skin, film grain, reflections, and material texture benefit from continuous detail modeling. Recent research is full of architectures designed to get both without forcing the whole system into one generative paradigm.
Which makes Chen’s refusal to choose “diffusion” or “autoregressive” look less evasive than it first appears. If the model can assign different noise states to different tokens, regions, or stages, there may be no single answer. The answer depends on what’s being generated, what’s already been fixed, and which part of the sampling schedule you’re looking at.
So instead of asking if GPT-Image-2.5 “is diffusion,” I want to know what gets diffused, in what order, and which parts of the image are able to stay clean. That’s the architecture question worth answering.
// solstice.eng
All posts →
