Luke Angel
Two ink-outlined picture frames side by side on a cream background. The left frame holds a clean row of evenly-spaced glyph blocks in olive-green; the right holds the same row rendered three times over, overlapping and misaligned, in muted grey. Faint dot grid, vertical olive-green accent bar at the left edge.

FLUX Can Spell, SDXL Cannot: Local AI Image Models on a DGX Spark

by
#ai#local-llm#dgx-spark#image-generation#flux#evals

Two models make pictures on this box: FLUX.1-schnell and SDXL-Turbo. Twenty-seven images later, the headline is one sentence.

FLUX can spell. SDXL cannot.

Everything else is detail — but the detail is where the interesting parts live, including two conclusions from an earlier round that turned out to be wrong about the models and right about my own code, which by now is the most predictable outcome in this project.

The method, which matters more than the pictures

Every probe carries three fields written in this order: the prompt, what to expect, and what we got. The expectation is written from the prompt alone, before generating anything.

This sounds like bureaucracy. It's the only thing standing between an honest result and a page of pretty pictures with captions retrofitted to match. Diffusion output is enormously persuasive at a glance, and a caption written after looking will always find something to praise. Writing the target first means the page can't quietly move it.

Every prompt is also chosen to break something. Counting, spatial relations and attribute binding are the three most reliable ways to embarrass a diffusion model, so the funny prompts and the hard prompts are the same prompts. A rubber duck at a cash machine is a counting test wearing a joke.

The headline: legible text

Three exact strings, scored character by character.

targetFLUX.1-schnellSDXL-Turbo
DGX SPARK REPAIRSDGX SPARK REPAIR — dropped the final SDEGXK PPARR- / SAX RR-ERK DGPARRRESS
HARNESS BUGS: 44exactHARNESSS HARKNESS / BUGS ARK 445 BAGS BAYS 4.54 44
THE ORCHESTRATORTHE ЛCHESTRATOR — the O collapsed, R missingTHIE THE ORECHESSETTOR SEALATOR

FLUX: one exact, one off by a character, one partial. SDXL: 0 for 3.

The same chalkboard prompt rendered by both models, side by side. FLUX.1-schnell produces a single lecture-hall blackboard reading HARNESS BUGS: 44 in neat white chalk — the colon, both words and both digits exactly as specified, with plausible pressure variation in the strokes. SDXL-Turbo produces a wider classroom shot whose blackboard carries the phrase written three times over and differently wrong each time: HARNESSS HARKNESS, BUGS ARK 445, BAGS BAYS 4.54, and finally a correct 44. The individual letter shapes are crisp in both images, so the failure is not blur — it is repetition.

The failure mode is the interesting part. SDXL's problem isn't blur — the letter shapes are individually crisp. Its signature failure is repetition: it renders the phrase two or three times, each version differently wrong. On the chalkboard it eventually wrote 44 — after also writing 445 and 4.54. It isn't failing to draw letters. It's failing to stop.

FLUX has a related tic. Asked for a book cover with a title and nothing else, it produced a clean product shot with a correct spine and shadow — and invented a two-line subtitle of pure gibberish. Unprompted text is where even the good model reverts to texture. That was in the written prediction, and it arrived on schedule.

What legibility costs, drawn as a comparison. SDXL-Turbo at 512 pixels loads in four and a half seconds and generates in under half a second, using about eight gigabytes. At 1024 it takes 1.2 seconds per image and 11.5 gigabytes, and scores zero out of three on legible text. FLUX.1-schnell at 1024 takes 217 seconds to load, 7.6 seconds per image, and 37 gigabytes of peak memory, scoring roughly two and a half out of three. The ratio is marked: 6.3 times the time and 3.2 times the memory. A note beneath states that this is not a trade-off when the picture contains words, because the cheaper option produces nothing usable.

6.3× the time and 3.2× the memory. Which sounds like a trade-off and isn't. If the picture contains words, SDXL is not the cheaper option — it isn't an option. Spend the 7.6 seconds or change the brief.

Resolution is a per-prompt decision, not a setting

The most useful practical finding here is one I nearly deleted.

SDXL-Turbo is a 512-native model. Pushed to 1024 it duplicates single subjects. Asked for "a lone lighthouse," it produced two — one on the cliff, one on a sea stack behind. The prompt contains the word lone, and the extra resolution overrode it.

The obvious lesson would be "run SDXL at 512." That lesson is wrong. On the same sweep, at the same 1024:

  • the overgrown data centre gained vines across the ceiling and green LED text on the racks — a clear win
  • the clockwork bird schematic became, arguably, the most detailed image produced in the whole set — from the model that loses to FLUX everywhere text is involved

The duplication artefact caught in the act. On the left, SDXL-Turbo at its native 512 pixels renders the prompt "a lone lighthouse on a basalt cliff at dusk" as exactly one lighthouse on a cliff above a storm sea. On the right, the identical model and prompt pushed to 1024 pixels renders two lighthouses — one on the cliff and a second on a sea stack behind it — despite the word "lone" in the prompt. Only the resolution setting changed between the two images.

Why resolution cannot be set globally. Two prompts run through the same model at the same 1024-pixel setting. The first names a single countable subject — a lone lighthouse — and the extra resolution tiles it, producing two lighthouses despite the word "lone" in the prompt. The second names no countable subject, and the same setting buys genuine additional detail: vines across a ceiling, readable indicator text on equipment racks. The deciding factor is marked as the presence of a countable noun in the prompt, not any property of the setting.

The difference is whether the prompt contains a countable subject. With one lighthouse to duplicate, 1024 breaks the image. With no single subject to tile, 1024 buys real detail. Same model, same setting, opposite outcomes, decided by the noun in the prompt.

The 512 set is the control that makes this visible at all. I deleted it during a rewrite for being redundant and was caught. Without it, "SDXL duplicates at 1024" is an assertion; with it, it's a comparison.

The hard set: six probes designed to fail

probeFLUX.1-schnellSDXL-Turbo
counting + text (5 ducks)partial — six ducksfail — ~19 ducks
spatial: on vs underpassfail — nothing under, nothing on
attribute binding (3 penguins)near — bow tie bled one positionfail — four penguins, all bow ties
hands + tools + signpartial — correct grips, three armsfail — two raccoons fused
style + long captionpassHERE ENDETH THE SPRINT exactfail — caption illegible
scale inversionpassfail — concept fusion

FLUX: 3 pass, 1 near, 2 partial, 0 fail. SDXL: 0 for 6.

Neither model can count. Asked for exactly five ducks, FLUX gave six and SDXL gave about nineteen. Diffusion models have no counting mechanism; they render "a queue of ducks" and the number falls out of the composition. Off by one versus off by a factor of four is a real difference in degree, but neither is a model you can ask for a specific number of things.

Attribute binding compared on the same three-penguin prompt. FLUX renders three penguins and binds two of the three requested accessories correctly — sunglasses on one, a business suit and briefcase on another — while the red bow tie bleeds one position to the right and appears on a penguin that should not have it. SDXL renders four penguins instead of three, puts an identical red bow tie on every one of them, and drops the sunglasses and the briefcase entirely. The most visually dominant attribute has won and spread across the whole image.

Attribute bleed is the cleanest divider. Given three penguins with three different accessories, FLUX bound two of three correctly and let the red bow tie bleed one position right. SDXL applied the bow tie to all four penguins — it miscounted too — and dropped the other two attributes entirely. The textbook failure in miniature versus the textbook failure at full strength: the most visually dominant attribute wins and spreads.

FLUX's failures are the predicted failures, arriving on time. It got both tools into the correct paws — the hard part — then grew a third arm to hold everything. In every case the written prediction named the failure before the image existed. A model that fails where you expect is far more useful than one that fails at random.

And the prettiest image of the entire run is an SDXL failure: a genuinely beautiful illuminated manuscript, gold leaf and all, featuring two knights, no rubber duck, and an illegible caption. It answers none of the prompt. Grading on beauty would have ranked it first.

What I'd tell a team

Write the expectation before you generate. For text output you can get away with judging after the fact, because a wrong answer usually looks wrong. Image output doesn't work that way — it looks good while being wrong, and your judgement adapts to whatever appeared.

And keep the control set even when it looks redundant. Mine was the difference between an assertion and a comparison, and I'd already deleted it once.

What's next

The same method, applied to video, where a model can render a completely convincing scene and simply decline to perform the action you asked for.

Keep reading

shares tags: #ai · #local-llm
tools
Nine Local LLMs Ranked by Cost Per Solved Task — Seven of Them Have No Price at All
Aug 17
method
1,200 Lines of Multi-Agent Orchestration, Beaten by One Local LLM Agent
Aug 10
tools
The Agent Framework Bake-Off: LangGraph vs Pydantic AI vs Hand-Rolled, and the 32 Lines That Mattered
Aug 07