Nano Banana 2 Lite is live on VibeArt: Gemini 3.1 Flash Lite Image costs 3 credits per 1K generation, so 100 trial credits cover roughly 33 fast image runs. Try Nano Banana 2 Lite
Qwen Image 3 Prompt Guide: 13 Real Tests for Text, Infographics, and UI
We test Qwen-Image-3.0 with 13 real prompts: dense text, multilingual 3×3 grids, long newspapers, nested UI, editing, and its exact failure modes.
July 26, 2026
•By VibeArt Team•
25 min read
Qwen-Image-3.0 is not interesting because it can make another attractive picture. It is interesting because Qwen is trying to make one image carry the structure of a deliverable: small copy, formulas, charts, interface hierarchy, regional layout, and knowledge.
That is a much harder promise to evaluate. A poster can look polished while spelling a name incorrectly. An app screen can contain every requested component while putting them in the wrong information hierarchy. A financial infographic can render a number clearly and still make the number false.
This guide therefore uses a stricter question than “does the demo look good?”:
Can Qwen Image 3 produce a useful visual document, or does it produce a beautiful screenshot that still needs line-by-line auditing?
Quick Answer: Is Qwen Image 3 Actually Useful?
Qwen-Image-3.0, released on July 21, 2026, is officially positioned for content-dense image generation: up to 4.5K tokens of input, text as small as roughly 10px, native rendering across 12 languages, complex document layouts, realistic detail, and interface simulation.
Our thirteen retained first-attempt outputs from the official Qwen Chat product show that the claim has substance. The ten baseline tests reproduced a dense 3×3 infographic, five mathematical expressions, six writing systems, a supplied financial dataset, and a newspaper front page with unusually high text accuracy. Three additional stress tests then pushed the official launch claims much harder: a 3,646-character multilingual 3×3 knowledge wall, a 3,812-character six-column newspaper, and a 3,538-character four-layer nested interface.
It is not deterministic layout software. The mobile UI added three stray slash glyphs and invented unreadable text inside its image preview. The wide scene delivered the requested three boats but produced five drones instead of four and three bridges instead of two. Editing changed the correct cup, preserved the sign, and kept the composition close, but reduced the output from 1792×2400 to 1072×1440 and introduced small non-target pixel changes.
The practical rule is simple:
use Qwen Image 3 when visual structure, information density, and mixed text-plus-image composition matter
provide exact source copy and source data instead of asking the model to invent facts
audit every character, number, formula, region, and UI relationship before delivery
keep facts and editable source material outside the generated bitmap
For mood boards and visual direction, small errors may be tolerable. For exams, scientific pages, pricing tables, financial charts, product UI, or public information, the generated image should be treated as a draft until it passes an explicit QA checklist.
What Qwen-Image-3.0 Is — And What Is Not Public Yet
The official launch describes the model with one word: “Real.” Qwen splits that into three areas:
Rich Content — long instructions and dense layouts such as newspapers, storyboards, exam papers, presentations, and nested interfaces.
Authentic Details — small text plus fine visual texture such as pores, hair, skin, and physical materials.
Deep Knowledge — multilingual rendering, familiar software and consumer interfaces, and knowledge-rich compositions.
The launch shows a complete 3×3 infographic generated in one pass rather than assembled from nine separate images. That is a meaningful product direction. It also raises the burden of proof: nine clean regions are not enough if one formula, date, label, or diagram is wrong.
As of July 25, 2026, we could verify and use Qwen-Image-3.0 in a signed-in official Qwen Chat account. A fresh image conversation defaulted to Qwen-Image 2.0, so we explicitly selected Qwen-Image 3.0 before every test and confirmed the visible model label after submission.
Alibaba Cloud now documents the API model ID qwen-image-3.0-pro, but it is in an invite-only preview, not general availability. The documented China-mainland limit is one request per minute, and the pricing page lists a limited-time free preview rather than a durable general-availability price. We still could not verify downloadable 3.0 weights, a public 3.0 model card, technical report, license, parameter count, or fixed maximum output resolution. The public Qwen-Image repository has not added a 3.0 release entry.
That distinction matters:
Item
Status we could verify on July 25, 2026
Official launch
Confirmed
Qwen Chat access
Confirmed in a signed-in account
Standalone 3.0 API model ID
qwen-image-3.0-pro
API availability
Invite-only preview
API price and rate limits
Limited-time free preview; 1 request per minute in China mainland
Downloadable 3.0 weights
Not confirmed
3.0 model card / technical report
Not confirmed
3.0 open-source license
Not confirmed
Maximum output resolution
Not confirmed
Do not infer that 3.0 is open source because earlier Qwen-Image releases were open. Do not infer a production API contract from the consumer chat product.
What Changed From Qwen-Image-2.0
Qwen-Image-2.0 was already positioned around professional infographics, unified generation and editing, native 2K output, roughly 1K-token instructions, and a lighter architecture. The 3.0 story moves from “better image generation with text” toward “generate the whole document-like visual in one pass.”
Knowledge-rich visual composition and familiar interfaces
Public deployment details
Generally documented API examples
qwen-image-3.0-pro invite-only preview
The useful upgrade is not merely “more prompt tokens.” It is the attempt to preserve hierarchy across many independent constraints. That is why our tests score structure and correctness separately.
How We Test: A Visual Document Needs More Than A Beauty Score
Every test records:
entry point and visible model label
generation date
requested and actual dimensions
generation time
total attempts and first usable attempt
exact-text errors
factual or calculation errors
layout and region errors
UI-semantic errors
non-target drift for editing
estimated manual repair time
We do not assign one overall score. A single number hides the failure that matters. Instead, each output is reviewed on seven dimensions:
Dimension
What we check
Exact text
Case, punctuation, digits, symbols, line breaks, and spelling
Factual accuracy
Dates, calculations, scientific statements, units, and supplied data
Layout adherence
Region count, order, relative position, missing and duplicated modules
UI semantics
Hierarchy, component relationships, touch targets, and familiar product logic
Edit preservation
Whether anything outside the requested edit changed
Visual quality
Material, lighting, anatomy, composition, and local artifacts
Workflow efficiency
Waiting time, retries, first-usable rate, and manual repair time
All outputs below were generated on July 25, 2026 in the signed-in official Qwen Chat web product. Each retained result is the first attempt after explicitly selecting and confirming the visible Qwen-Image 3.0 label. Queue time is approximate because the consumer UI exposes no machine-readable timing. Qwen does not expose a prompt-expansion switch in this interface, so we cannot claim that internal prompt rewriting was disabled.
Tests 1–10 use downloaded outputs converted to quality-90 WebP without resizing. Tests 11–13 use optimized WebP copies without resizing. All images are served from our CDN at their native pixel dimensions. The complete prompt used for every test appears below its result; the three longest prompts are collapsed by default to keep the article readable.
Test 1: Dense 3×3 Infographic And Exact Copy
This test separates layout success from content success. It contains nine regions, 27 required bullet points, formulas, status codes, and exact labels.
Result: 2048×2048, approximately 173 seconds, one attempt. All nine regions appeared in the requested order. The headings, scientific statements, status codes, and color equations are legible and correct. Two character-level deviations prevent a literal-perfect score: the compound-interest expression was typeset as a real superscript instead of preserving the source characters ^(nt), and the binary-search diagram contains an extra R label. Estimated manual repair: 1–2 minutes in an editable layout, but the bitmap itself is not conveniently repairable.
Prompt:
Create a square 1:1 educational infographic with a 3×3 grid on off-white. Each cell has a navy title bar, one simple diagram, and only the specified copy. Do not add, remove, paraphrase, translate, or repeat text.1 WATER CYCLE: Evaporation—liquid becomes vapor; Condensation—vapor forms clouds; Precipitation—water returns to Earth.2 PHOTOSYNTHESIS: Inputs—water, carbon dioxide, sunlight; Output—glucose and oxygen; Location—chloroplasts.3 NEWTON'S LAWS: First—motion stays unchanged without net force; Second—F = ma; Third—every action has an equal opposite reaction.4 BINARY SEARCH: Requires a sorted list; Compare with the middle value; Time complexity—O(log n).5 COMPOUND INTEREST: Formula—A = P(1 + r/n)^(nt); P is the principal; A is the final amount.6 DNA: Bases—A, T, C, G; Shape—double helix; A pairs with T; C pairs with G.7 SUPPLY AND DEMAND: Higher price can reduce demand; Higher price can increase supply; Equilibrium is where curves meet.8 HTTP STATUS: 200 OK; 404 Not Found; 500 Internal Server Error.9 COLOR MIXING: Red + blue = magenta; Blue + green = cyan; Red + green = yellow.Use consistent editorial styling. All text legible. No logo, watermark, or footer.
Test 2: Small Text, Formulas, And An Academic Page
A formula that merely resembles the source is a failure. We check primes, limits, integral bounds, superscripts, minus signs, and every digit.
Result: 1792×2400, approximately 89 seconds, one attempt. All five mathematical statements are correct, including the limit, derivative, exponents, integral bounds, prime mark, logarithm, and fractions. The title, date, and footer are correct. The requested underscore characters after Name: became a conventional horizontal rule, so the mathematics passes while strict character preservation does not. Estimated manual repair: under 1 minute if the source remains editable.
Prompt:
Create a portrait A4 mathematics practice sheet, photographed flat from directly above, black text on white paper.Exact title: "Limits and Derivatives — Practice Sheet"Exact subtitle: "Name: ____________ Date: 2026-07-25"Include exactly these five numbered items:1. lim(x→0) sin(x)/x = 12. d/dx [x³ + 2x² − 5x + 7] = 3x² + 4x − 53. f(x) = e^(2x), therefore f′(x) = 2e^(2x)4. ∫₀¹ 3x² dx = 15. If y = ln(x² + 1), then dy/dx = 2x/(x² + 1)Footer exactly: "Check every exponent, limit, fraction, and derivative symbol before submitting."No handwriting, extra formulas, logo, or watermark.
Test 3: Multilingual Text Rendering
This test uses six writing systems and gives Arabic equal prominence. We verify exact spelling, character mixing, direction, punctuation, and whether the model quietly adds decorative copy.
Result: 1696×2528, approximately 112 seconds, one attempt. English, Simplified Chinese, Japanese, Korean, Spanish, and Arabic all match the supplied strings. Arabic is rendered right-to-left, and no seventh line or decorative pseudo-text appeared. Visual size is not perfectly equal—the English line is wider—but the hierarchy does not demote a language. Estimated manual repair: none.
Prompt:
Create a portrait international bookstore poster with a dark green background, cream paper texture, and one central open book. It must contain exactly these six lines and no other text:Stories Across Borders跨越边界的故事国境を越える物語국경을 넘는 이야기Historias sin fronterasحكايات بلا حدودOne language per line, exact spelling, Arabic right-to-left, equal visual importance. No logo, watermark, or invented characters.
Test 4: Realistic UI Versus Correct UI Logic
An interface can look plausible while behaving like no real product. Here, the critical constraint is that each comment author must appear above the comment text.
Result: 1536×2752, approximately 157 seconds, one attempt. The post hierarchy, three authors, comment order, action counts, and bottom navigation are all correct. Each author appears above the matching comment. It still fails copy cleanliness: a stray backslash-like glyph appears before every comment, and the image preview contains invented pseudo-CJK lettering. This is a convincing UI mockup, not a production screenshot. Estimated manual repair: 3–5 minutes in a design file.
Prompt:
Create a portrait 390×844 dark-mode mobile social app screenshot.Top: avatar, centered "Following", search and settings.Post hierarchy: "Maya Chen", "@mayac", "2h", text "Testing a multilingual poster workflow today.", one landscape preview, then reply 12, repost 28, like 316, bookmark.Exactly three comments, username ABOVE comment text:Leo Park — "The spacing feels much better now."Sofia Ruiz — "Can you test Spanish text next?"Amir Haddad — "Please include Arabic direction too."Bottom navigation: Home, Search, Create, Notifications, Profile; Home active. Believable spacing and touch targets. No extra comments, numbers, device frame, logo, or watermark.
Test 5: Exact Object Count And Spatial Constraints
Object-layout prompts are easy to audit. “Almost eight objects” is not a subjective result.
Result: 2048×2048, approximately 88 seconds, one attempt. Exactly eight object categories appear, every item is fully visible, and all eight occupy their assigned region. There are no duplicate categories or decorative extras. The passport is intentionally plain because the prompt also forbids text. Estimated manual repair: none.
Prompt:
Create a square top-down product photograph on a matte charcoal table. Show exactly eight objects and nothing else:1 red passport at top-left2 silver fountain pen at top-center3 black compact camera at top-right4 green ceramic cup at middle-left5 closed blue notebook at exact center6 white earbuds case at middle-right7 yellow keychain at bottom-left8 pair of round brown sunglasses at bottom-rightEvery object fully visible, separated, only in its assigned position. Soft studio shadows, realistic materials. No text, logo, watermark, or duplicates.
Test 6: A Financial Infographic With Supplied Facts
The model receives every number. We recalculate the total and percentage ourselves, then check whether visual polish changed the data.
Result: 2752×1536, approximately 138 seconds, one attempt. Every supplied value and label is correct: 12.4, 15.1, 14.5, 42.0, 45.0, 3.0, and 93.3%. The bars follow April–May–June order, May is tallest, the total sums correctly, and no extra metric was invented. Estimated manual repair: none.
Prompt:
Create a landscape financial infographic using only this data.Title: "Northwind Q2 Revenue"Subtitle: "All values in USD millions"April 12.4; May 15.1; June 14.5; Quarter total 42.0; Quarter target 45.0; Gap to target 3.0.Include one April-May-June bar chart, progress "42.0 of 45.0", callout "May was the strongest month", and callout "Target attainment: 93.3%". Every number, label, bar order, and calculation must match exactly. No other metrics, conversion, logo, or watermark.
Test 7: Photorealism, Age, Hands, And Material Detail
This is not a beauty portrait. The source scene also becomes the input for the preservation test that follows.
Result: 1792×2400, approximately 78 seconds, one attempt. The result has believable age detail, window light from the left, natural hands with five fingers each, clay residue, material texture, a white cup on the right, and the exact card text FIRE TO 1220°C. The composition became portrait-oriented even though no aspect ratio was specified; that is a product choice rather than a prompt violation. Estimated manual repair: none.
Prompt:
Documentary-style environmental portrait of a 58-year-old female ceramic artist seated at a wooden worktable in her studio. Natural window light from the left. Visible age lines, realistic pores, a few gray hairs, clay dust on both hands, small wrinkles in a linen apron, unfinished stoneware bowls on wooden shelves, subtle fingerprints on wet clay.Place one plain white ceramic cup on the right side of the table. Place a small paper studio card behind it containing exactly the text "FIRE TO 1220°C".Neutral color grading, 50mm lens look, shallow but believable depth of field. Preserve natural asymmetry and real skin texture. No glamour retouching, no waxy skin, no extra fingers, no other text, no logo, no watermark.
Test 8: Image Editing And Non-Target Preservation
The result from Test 7 is used as the only input. The requested edit is intentionally small; every other change counts as drift.
Original: 1792×2400Edited: 1072×1440
Result: approximately 61 seconds, one attempt. The cup becomes deep cobalt blue with one gold crescent, while the person, hands, apron, light direction, shelves, table, clay, and exact firing card remain visually close. This is not a pixel-local edit: the service reduces both dimensions by about 40%, and after resizing the source to the edited dimensions, the mean absolute RGB-channel difference outside a conservative cup mask is 9.4 on a 0–255 scale. The edit is visually useful, but resolution and non-target preservation must be checked before chaining more edits. Estimated manual repair: none for this request; re-upscaling or source compositing may be needed in production.
Prompt:
Edit only the ceramic cup on the right side of the table: change its color from white to deep cobalt blue and add a small gold crescent on its front.Preserve every other detail as closely as possible, including the person's identity, face, hands, clothing, lighting, camera angle, table objects, background shelves, and the exact existing text "FIRE TO 1220°C". Do not move, resize, add, or remove anything else.
Test 9: Newspaper Layout And Long Copy
This test checks columns, hierarchy, exact copy, date handling, and whether the model invents extra headlines to make the page look more complete.
Result: 1792×2400, approximately 141 seconds, one attempt. The masthead, date, main headline, deck, three checklist lines, weather copy, and footer all match. The output adds no invented article text and maintains a credible front-page hierarchy around the requested photograph. Estimated manual repair: none.
Prompt:
Create a portrait broadsheet newspaper front page photographed flat overhead. Use exactly this readable copy and no other text.Masthead: "THE DAILY WORKBENCH"Date: "Saturday, July 25, 2026"Main headline: "Small Teams Turn Visual Drafts Into Deliverables"Deck: "A new workflow keeps prompts, source data, revisions, and approvals connected."Left headline: "Three Checks Before Publishing"Bullets: "Verify every number"; "Compare every label"; "Preserve the editable source"Right headline: "Weather"Copy: "Morning cloud, afternoon sun, high 24°C."Footer: "Edition 01 • Price $2.00"Black, warm white, muted red; one photo of a designer reviewing a proof; believable hierarchy and columns. No invented copy, logo, or watermark.
Test 10: Extreme Wide Scene And Local Detail
Wide scenes expose repeated textures, weak anatomy, and a gap between nominal resolution and useful detail.
Result: 2752×1536, approximately 96 seconds, one attempt. The scene has three clearly different regions, exactly three solar boats, an airship, varied towers, more than 20 people, and recognizable walking, cycling, exercise, painting, picnicking, and chess activities. Count adherence fails: there are five visible delivery drones rather than four, and three bridge structures rather than two. Small human anatomy is acceptable at page size but soft at native zoom. Estimated manual repair: 5–10 minutes with object removal or regeneration.
Prompt:
Create an ultra-wide 16:9 realistic future ecological city.Left third: residential towers with varied vertical gardens and maintenance walkways, no repeated facade.Center: river with exactly three solar passenger boats and two pedestrian bridges.Right: public park with at least 20 visible people doing distinct activities including walking, reading, cycling, chess, and exercise.Sky: exactly four delivery drones and one transparent passenger airship.Coherent daylight, scale, anatomy, natural materials, continuous city plan. No text, logo, watermark, cloned people or trees, repeated windows, malformed limbs, or impossible bridges.
Extreme Stress Tests: Long Prompts, Dense Text, And Nested UI
The official launch does more than show attractive samples. Its harder examples combine a long prompt with either horizontal breadth—a complete 3×3 knowledge wall—or vertical depth, where one interface contains another interface that contains more structured content. We designed three new tests around those two axes.
These are deliberately unreasonable briefs for a normal image generator. The goal is not to prove that a bitmap should replace a newspaper layout tool, a localization workflow, or a code editor. The goal is to locate the boundary between structurally complete and literally correct.
Test
Prompt size
Structural result
Literal failure
11. Extreme 3×3 knowledge wall
3,646 characters
All nine domains present in a strict grid
Several smallest non-Latin labels degraded
12. Six-column broadsheet
3,812 characters
Every major newspaper module present
Lead copy and some Chinese copy were not exact
13. Four-layer nested UI
3,538 characters
All four interface layers and six runbook steps present
Code looked plausible but contained syntax and spelling errors
Test 11: Extreme 3×3 Multilingual Knowledge Wall
This is a substantially harder version of Test 1. Nine cells must cover incident command, spatial geometry, genomics, zero-trust networking, typhoon response, banking controls, 12-language wayfinding, a laboratory SOP, and ship-readiness checks. The prompt mixes formulas, tables, arrows, coordinates, status labels, and 12 scripts while forbidding omitted or invented modules.
Result: 2048×2048, approximately 130 seconds, one retained Qwen-Image 3.0 attempt. The model produced a strict 3×3 grid with all nine requested domains and a coherent visual system. English and Chinese headings, formulas, coordinates, most numbers, and the larger table entries remain surprisingly legible. The result reads as one designed artifact rather than nine unrelated thumbnails.
The failure appears exactly where the official small-text and multilingual claims become hardest to combine. Several micro-sized Russian, Arabic, and Hindi strings are malformed; small typhoon-legend copy and a few network-arrow labels collapse into pseudo-text. Some decimal punctuation is localized from a period to a comma. This passes structural completeness but fails literal multilingual QA.
View the exact 3,646-character prompt used for Test 11
请用 Qwen-Image 3.0 一次性生成一张超高信息密度、正方形 1:1、清晰可读的专业“3×3 应用知识图谱”海报。它不是九张图拼接,而是一张统一设计系统中的完整九宫格;严格 3 行×3 列,共 9 格,不多不少。画布四周留 64px 外边距,格间留 28px 白色沟槽,每格都有细灰边框、独立标题栏、编号 01–09、图表区、正文区和页脚来源标签。全图使用瑞士国际主义信息设计:白底、深墨蓝文字、青色/橙色/红色作为状态色,网格严谨,印刷级排版。主标题精确写:“QWEN IMAGE 3 — EXTREME 3×3 KNOWLEDGE WALL”;副标题精确写:“Nine domains · one canvas · zero omitted labels”。所有引号中的文字必须逐字呈现,不得改写、翻译、漏字或生成乱码;正文允许 10–14px 视觉字号但仍需清晰。严禁水印、品牌误拼、额外面板、跨格元素、重复图标。第 01 格,标题精确写:“INCIDENT COMMAND / 事故指挥”。左侧画一条 5 节点时间轴,逐行精确写:“09:17 API latency > 2.0s”“09:21 Error rate reaches 8.4%”“09:26 Read replica isolated”“09:34 Cache warmed to 92%”“09:41 Service fully recovered”。右侧是四张 KPI 小卡:“P95 684 ms”“Errors 0.7%”“RPS 12,480”“MTTR 24 min”。底部红框写:“DECISION: HOLD DEPLOYMENT”。第 02 格,标题精确写:“SPATIAL GEOMETRY / 空间几何”。绘制透明立方体 ABCD–A₁B₁C₁D₁、对角线 AC₁、平面 BB₁D₁D 和 35° 角标。右侧严格排出三行公式:“|AC₁| = √(a²+b²+c²)”“cos θ = (u·v)/(|u||v|)”“V = |a·(b×c)|”。下方定理框逐字写:“If u·n = 0, then u ∥ plane Π.” 和中文:“向量与法向量垂直,则向量平行于平面。”第 03 格,标题精确写:“GENOMICS PIPELINE / 基因组流程”。从左到右画五步流程:“Sample → Extraction → Library → Sequencing → Variant Call”。画双螺旋、FASTQ 文件、覆盖度直方图、染色体 7 的局部放大。小表格三列标题:“Variant”“Depth”“Effect”,三行精确数据:“chr7:140453136 A>T | 812× | missense”“chr12:25398284 C>G | 406× | intronic”“chr17:7674220 G>A | 1,024× | stop gained”。页脚写:“Reference: GRCh38 · QC PASS”。第 04 格,标题精确写:“ZERO-TRUST NETWORK / 零信任网络”。画清晰架构图:左侧“USER DEVICE”经过“IDENTITY PROXY”进入中心“POLICY ENGINE”,再分流到“WEB APP”“POSTGRES”“OBJECT STORE”;每条箭头旁分别标注“mTLS”“OIDC”“short-lived token”“deny by default”。右下角放 2×2 状态卡:“Auth 99.99%”“Blocked 1,284”“Risk HIGH 03”“Key age 17d”。底部精确写:“Never trust. Always verify. Log every decision.”第 05 格,标题精确写:“TYPHOON BRIEFING / 台风简报”。画东亚沿海地图、台风路径 6 个时间点、同心风圈、颜色图例。路径标签精确写:“T+00 18.2°N 126.4°E”“T+12 19.7°N 124.8°E”“T+24 21.3°N 123.1°E”“T+36 23.8°N 121.6°E”“T+48 26.1°N 120.9°E”“T+60 28.4°N 121.7°E”。右侧风险表:“Wind 145 km/h — SEVERE”“Rain 280 mm — EXTREME”“Surge 2.4 m — HIGH”。警示条写:“EVACUATE LOW-LYING AREAS BEFORE 18:00”。第 06 格,标题精确写:“BANK CONTROL MATRIX / 银行内控矩阵”。画 4×4 表格,列标题:“Control”“Owner”“Frequency”“Evidence”;四行精确写:“Dual approval | Treasury | Daily | Signed ledger”“Limit review | Risk | Weekly | Exception log”“Access recertification | Security | Quarterly | IAM report”“Backup restore | SRE | Monthly | Recovery record”。右侧画环形图:“PASS 87%”“OPEN 9%”“OVERDUE 4%”。底部橙框精确写:“3 HIGH-RISK EXCEPTIONS REQUIRE ACTION”。第 07 格,标题精确写:“12-LANGUAGE WAYFINDING / 十二语导视”。画现代地铁换乘站平面图,中央圆形“PLATFORM 4”,八条彩色线路,四个出口。右侧必须逐行准确出现以下 12 行,不合并:“English — Central Station”“中文 — 中央车站”“日本語 — 中央駅”“한국어 — 중앙역”“Español — Estación Central”“Français — Gare Centrale”“Deutsch — Hauptbahnhof”“Italiano — Stazione Centrale”“Português — Estação Central”“Русский — Центральный вокзал”“العربية — المحطة المركزية”“हिन्दी — केंद्रीय स्टेशन”。底部写:“Exit C · Lift · Tickets · First Aid”。第 08 格,标题精确写:“LAB SOP / 实验室规程”。画分子结构、烧杯、移液器、离心机和四枚标准危险图标。右侧编号步骤逐字写:“1. Wear nitrile gloves and goggles.”“2. Add 25.0 mL buffer at 4°C.”“3. Centrifuge 12,000×g for 10 min.”“4. Transfer exactly 200 μL supernatant.”“5. Record batch ID before disposal.” 参数卡写:“pH 7.40 ± 0.05”“Temp 4°C”“RPM 13,200”“Timer 00:10:00”。红色页脚:“DO NOT MIX ACID WITH HYPOCHLORITE”。第 09 格,标题精确写:“SHIP-READINESS CHECK / 发布检查”。画 12 项双列核对清单,每项前有绿色勾或橙色圆点,文字精确为:“Schema migration reviewed”“Rollback tested”“Feature flag off by default”“Secrets rotated”“Error budget checked”“Accessibility AA”“Mobile 360px verified”“P95 under 800 ms”“Audit log retained”“Runbook linked”“On-call acknowledged”“Customer notice drafted”。右下角盖章精确写:“GO WITH CONDITIONS”。页脚版本号:“REV 3.0 · 2026-07-25 · OWNER: PLATFORM”。整张图必须在一张画布中维持九格互不干扰:每格视觉语言统一但内容图形各异;所有表格行列对齐;数字、标点、上下标、希腊字母、中文、阿拉伯文、天城文都清楚;无随机伪文字,无文本块缺失。优先保证文字准确性、九格完整性和可判读性,其次才是装饰。
Test 12: Six-Column Newspaper With Long-Form Copy
The baseline newspaper used a small set of exact headlines and short copy. This version asks for an authentic six-column broadsheet with a masthead, date and price line, a lead story titled LUNAR WATER TREATY CLEARS FINAL VOTE, three supplied English paragraphs, a lunar photograph and caption, KPI cards, a line chart, a five-row market table, four dispatches, a Chinese-language section, a research formula and data table, a briefing box, and a literal footer.
Result: 1792×2400, approximately 75 seconds, one attempt. The hierarchy is excellent. Every major module is present, the masthead and primary headline are correct, the KPI and market table are readable, and the chart, dispatches, Chinese section, equation, stratigraphy diagram, research table, briefing, email address, and page number all occupy believable newspaper positions.
It is not copy-perfect. The opening English paragraph contains overlapping and corrupted words, some Chinese body copy and a byline diverge from the supplied source, and several long sentences are paraphrased or shortened. The image is an unusually convincing editorial comp, but publishing it without typesetting the copy again would be unsafe.
View the exact 3,812-character prompt used for Test 12
请使用 Qwen-Image 3.0 生成一张竖版 4:5、2048px 级别、可直接印刷的真实英文数据报纸首页。它不是海报,而是一张完整 broadsheet newspaper front page:纸张有轻微米白纤维、墨色细微渗透、折痕和套印误差,但所有文字必须清晰。采用严格 6 栏网格、细黑分隔线、经典衬线正文字体、无衬线数据标签。主报头逐字写:“THE ORBITAL LEDGER”;其下细字逐字写:“Independent reporting for Earth, Moon and Mars”;左上日期:“FRIDAY, JULY 25, 2026”;右上:“VOL. 18 · NO. 207 · PRICE $3.50”。不得生成伪文字,不得用无意义 lorem ipsum,不得遗漏引号中的任何段落。所有正文保持真实报纸密度,最小字看起来约 10px 但放大可读。第一、二、三栏组成头版主稿。超大标题精确写:“LUNAR WATER TREATY CLEARS FINAL VOTE”。副标题精确写:“Twelve nations agree on shared extraction limits, emergency reserves and open telemetry after a 19-hour session.” 署名精确写:“BY MAYA CHEN · SCIENCE & POLICY EDITOR”。正文必须逐字排出以下三段,不改写:“GENEVA — Delegates approved the Lunar Water Accord at 04:12 local time on Friday, creating the first enforceable limits on polar ice extraction. The agreement assigns annual quotas, protects permanently shadowed craters, and requires every operator to publish machine-readable telemetry within sixty seconds.”“Negotiators resolved the final dispute by placing a six-month emergency reserve under joint control. Any release now requires three independent signatures: one from the host nation, one from the scientific council, and one from the civilian settlement network.”“Markets reacted cautiously. Oxygen futures fell 2.8 percent while launch-insurance shares rose. Engineers welcomed the data clause but warned that sensor calibration standards must be finished before the treaty takes effect on January 1, 2027.”主稿中间放一张写实月球南极基地照片:低角度阳光、机械臂、冰样品罐、三名宇航员;照片说明逐字写:“Survey team 7A returns from Shackleton Ridge with sealed core samples. PHOTO: LUNA PRESS POOL”。第四栏顶部做数据模块,标题:“THE NUMBERS”。四行 KPI 精确写:“12 signatory nations”“18.6 million liters annual cap”“15% emergency reserve”“60 seconds telemetry delay”。其下画折线图标题:“WATER OUTPUT, 2027–2032”,横轴年份 2027、2028、2029、2030、2031、2032,纵轴标注“million liters”,三条线图例严格写:“Approved quota”“Projected demand”“Protected reserve”。图下注释:“Forecast range: ±6.2%”。第五栏顶部标题:“MARKET SIGNALS”。画 5 行金融表格,列标题严格为:“Asset”“Close”“Change”。数据逐行精确写:“OXY-FUT 84.20 −2.8%”“HELIO 118.44 +1.6%”“CISLUNAR 72.09 +0.4%”“ICE-INDEX 204.81 −1.1%”“LAUNCH-RE 51.77 +3.2%”。表下短文逐字写:“Volume was 31 percent above the 20-day average. Analysts cited lower scarcity risk and stronger compliance spending.”第六栏顶部标题:“FIELD DISPATCHES”。列出四条带地点的小稿,每条标题和句子都要准确:“SHACKLETON — Power restored” 下写:“A damaged 40 kV cable was isolated at 06:30; habitats remained on battery for eleven minutes.”“MARE IMBRIUM — Dust alert” 下写:“Electrostatic dust reached level orange. Exterior maintenance is paused until 14:00 UTC.”“TYCHO — School opens” 下写:“The first bilingual primary school welcomed 86 students from nine settlement zones.”“L1 GATEWAY — Docking delay” 下写:“Cargo vehicle Kestrel-4 will berth six hours late after a guidance-software rollback.”页面下半部横跨前三栏做中文栏目。栏目名准确写:“中文速览”。标题:“月球水资源公约通过最终表决”。正文两段逐字写:“十二个参与国同意公开开采遥测数据,并为科研、民用与应急储备设定独立配额。公约将永久阴影区列为重点保护区域,任何商业作业都必须保留完整审计记录。”“联合执行委员会每季度发布一次透明度报告。若连续两次未能上报数据,相关运营方将自动暂停许可证,直至第三方完成安全复核。”右下角署名:“记者 陈雨 · 日内瓦”。页面下半部中间做科研插页,标题:“RESEARCH NOTE: ICE STABILITY”。准确排出公式:“∂C/∂t = D∇²C − kC” 和 “T_eq = 112 ± 4 K”。画三层剖面:“REGOLITH 0–18 cm”“MIXED ICE 18–64 cm”“DENSE ICE >64 cm”。旁边小表格列标题:“Depth”“Purity”“Confidence”,三行:“22 cm | 31% | 0.82”“48 cm | 67% | 0.94”“81 cm | 91% | 0.97”。注释逐字写:“n = 214 cores · instrument drift corrected”。页面最下方横贯六栏的短讯带标题:“BRIEFING”。依次排出四条:“Solar weather: G2 storm watch begins 16:00 UTC.” “Earth desk: Pacific heatwave breaks 14 records.” “Mars desk: Valles Marineris relay returns to service.” “Culture: The zero-gravity orchestra announces an autumn tour.” 最底部页脚精确写:“orbitalledger.news · Corrections: corrections@orbitalledger.news · Printed with 62% recycled fiber”。右下角印页码:“A1”。整体要求:像真实获奖报纸设计,不像 AI 海报;六栏从头到尾对齐;长段落完整;英文大小写、连字符、百分号、邮箱、数学符号、中文标点准确;照片与图表服务于内容但不遮挡文字;绝无随机字母串、重复段落、水印或多余标题。优先保证所有指定文本逐字可读。
Test 13: Four-Layer Nested Software Interface
This test asks one image to preserve semantic hierarchy across four visible layers: a VS Code-style desktop shell; an embedded browser showing an Atlas Ops incident dashboard; a phone chat overlay; and a runbook card nested inside that phone UI. The brief also supplies source code, terminal logs, KPIs, a network map, a graph, an incident table, chat messages, and six ordered runbook steps.
Result: 2752×1536, approximately 125 seconds, one attempt. All four layers are present and their boundaries are immediately understandable. The dashboard retains the map, graph, KPI hierarchy, timeline, and incident table. The phone contains the requested conversation, and the innermost runbook preserves all six steps plus owner, rollback, check, and rule fields. This is a strong demonstration of UI composition and semantic nesting.
The code editor is the weak point. Function names and paths drift, function is misspelled, braces and line structure are inconsistent, and some lines are duplicated. A decimal value also changes punctuation. The result is excellent as a product concept image and dangerous as a literal code screenshot: plausible syntax is not valid syntax.
View the exact 3,538-character prompt used for Test 13
请用 Qwen-Image 3.0 一次生成一张 16:9、2560×1440 视觉分辨率、极清晰的“四层嵌套软件界面”真实产品截图。核心测试是纵向语义深度:外层是 VS Code;VS Code 的右侧预览里嵌套一个浏览器运维仪表盘;仪表盘中弹出一个手机协作聊天窗口;聊天里又嵌入一张微型事故处置卡片。四层必须同时完整、边界清晰、透视一致,不能把某层的按钮或文字串到另一层。所有引号里的文本必须逐字呈现,代码标点、日志时间、表格数字、手机消息和最内层卡片都要可读。不要水印,不要随机伪文字,不要重复窗口。第 1 层:完整 VS Code 深色界面,占整张图。顶部 macOS 标题栏和交通灯按钮,窗口标题精确写:“atlas-ops — Visual Studio Code”。菜单精确写:“File Edit Selection View Go Run Terminal Help”。左侧 Activity Bar 有 Explorer、Search、Source Control、Run、Extensions 五个标准图标。Explorer 标题:“ATLAS-OPS”,文件树准确显示:“src”“app”“api”“incidents”“route.ts”“components”“IncidentPanel.tsx”“lib”“telemetry.ts”“package.json”“README.md”。打开的标签页有:“route.ts”“IncidentPanel.tsx”“telemetry.ts”。底部状态栏逐项写:“main*”“0 errors”“2 warnings”“Ln 18, Col 27”“Spaces: 2”“UTF-8”“TypeScript”“Prettier”。VS Code 中央编辑区打开 src/app/api/incidents/route.ts,显示带语法高亮、行号 1–18 的 TypeScript 代码,代码必须尽量准确排出:“import { NextResponse } from 'next/server'”“import { getIncidentSummary } from '@/lib/telemetry'”“”“export async function GET() {”“ const summary = await getIncidentSummary({”“ service: 'payments',”“ windowMinutes: 30,”“ includeRegions: ['us-east', 'eu-west', 'ap-south'],”“ })”“”“ return NextResponse.json({”“ generatedAt: '2026-07-25T09:42:18Z',”“ status: summary.errorRate > 0.02 ? 'degraded' : 'healthy',”“ summary,”“ })”“}”当前行高亮在第 13 行。左下 Terminal 面板标题精确写:“TERMINAL OUTPUT DEBUG CONSOLE PORTS”。终端命令与输出逐行写:“$ npm run dev”“▲ Next.js 16.2.6”“✓ Ready in 842ms”“GET /api/incidents 200 in 47ms”“WARN payments error_rate=0.034 region=eu-west”“INFO fallback_route=active recovered=92%”。第 2 层:VS Code 右侧约 44% 宽区域是内嵌浏览器预览,但仍在 VS Code 编辑器分栏内。浏览器地址栏精确写:“https://ops.atlas.example/incidents/INC-2048”。页面是浅色专业运维仪表盘,页眉左侧标志:“ATLAS OPS”,面包屑:“Incidents / INC-2048”,右侧绿色状态:“LIVE · updated 09:42:18 UTC”。主标题:“Payment authorization latency”。红色标签:“SEV-1 ACTIVE”。副标题:“Elevated p95 and issuer timeouts in EU West”。四张指标卡逐字写:“P95 LATENCY 2.84 s”“ERROR RATE 3.4%”“AFFECTED 18,240”“RECOVERED 92%”。仪表盘左中画欧洲地图,三枚节点标签:“DUBLIN 4.1%”“FRANKFURT 3.7%”“PARIS 2.9%”,红色链路从 Dublin 指向 Frankfurt。右中折线图标题:“Latency by region · last 30 min”,图例:“eu-west”“us-east”“ap-south”,横轴从 09:12 到 09:42,红线在 09:31 峰值 3.2s。下方时间线表格列标题:“Time”“Event”“Owner”“State”,四行精确写:“09:17 | Alert fired | SRE Bot | OPEN”“09:23 | Traffic shifted 20% | Mina K. | DONE”“09:31 | Issuer timeout spike | Payments | WATCH”“09:39 | Cache warm 92% | Leo R. | DONE”。右下按钮:“OPEN RUNBOOK”“ACKNOWLEDGE”“START BRIDGE”。第 3 层:仪表盘右下角覆盖一个约页面高度 48% 的手机协作窗口,外形像真实现代手机但完全在浏览器预览内部,不可溢出到 VS Code。手机顶部频道名:“# incident-inc-2048”,旁边小字:“8 members · bridge live”。消息按时间排列,头像与气泡清晰:“09:24 Mina — Shifted 20% of EU traffic. Error rate is falling.”“09:28 Leo — Cache warm-up reached 74%. ETA six minutes.”“09:33 SRE Bot — Threshold breached: issuer_timeout > 3.0%.”“09:37 Priya — No data loss. Retries remain idempotent.”“09:41 Mina — Recovery at 92%. Holding deployment until 10:15.”输入框占位文字:“Message #incident-inc-2048”,右侧按钮:“Send”。顶部红点提示:“1 unresolved decision”。第 4 层:在手机里 Priya 的消息下方嵌一张小而完整的事故处置卡片,像聊天附件,卡片标题精确写:“RUNBOOK R-17 · PAYMENT DEGRADATION”。橙色状态条:“CURRENT STEP 4 OF 6”。六步纵向流程必须出现并带编号:“1 Confirm telemetry”“2 Freeze deploys”“3 Shift 20% traffic”“4 Warm issuer cache”“5 Validate retries”“6 Restore gradually”。第 1–3 步绿色打勾,第 4 步橙色高亮,第 5–6 步灰色。卡片底部三项细字:“Owner: Payments SRE”“Rollback: ready”“Next check: 09:46 UTC”。最下方红色规则:“STOP IF ERROR RATE > 5%”。整体要求:外层 VS Code 像真正桌面应用,内层浏览器像真正 SaaS,第三层手机像真实协作工具,第四层卡片像真实附件;四层尺度递减但每层关键文字仍清晰;每一层只显示属于自己的导航和控件;代码、终端、指标、地图、折线图、时间线、手机消息、六步 runbook 全部存在。优先级依次是:层级完整 > 文本准确 > UI 真实性 > 装饰。
Together, these three outputs expose the model's current ceiling. Qwen-Image-3.0 is unusually good at preserving many modules and their relationships in one pass. Exactness degrades faster than structure as text gets smaller, the number of scripts grows, or the copy becomes code-like. That distinction should determine whether the output can ship or must remain a visual draft.
Where Qwen Image 3 Still Fails
The failure categories matter more than a leaderboard:
Clear text can still be wrong text. Readability is not exactness.
A correct layout can contain incorrect knowledge. The diagram and the labels require separate checks.
A realistic interface can violate interface logic. Review hierarchy and relationships, not just pixels.
One-pass density increases audit cost. More content per image means more opportunities for a silent error.
A bitmap is not an editable document. Preserve source copy, data, and layout intent outside the generated output.
A successful edit can still drift elsewhere. Compare identity, copy, lighting, and composition against the input.
Multilingual accuracy falls with character size. A model can pass a six-line language poster and fail the same scripts as microcopy inside a dense grid.
Plausible code is not executable code. Treat generated editor screenshots as visual concepts unless every token is independently verified.
An early third-party comparison published after launch reported the same important distinction: Qwen could produce convincing layout and realistic scenes while still making an Arabic-text error, a knowledge error in a science poster, and an information-hierarchy error in a social UI. We treat those observations as test hypotheses, not as our own results.
A Reusable Qwen Image 3 Prompt Formula
Long prompts work better when they behave like production briefs:
[deliverable type][canvas / aspect ratio / visual hierarchy][exact source copy or source data][region-by-region composition][style, material, and lighting][non-negotiable constraints][explicit forbidden additions][verification instruction]
Five practical rules:
provide exact source copy; do not ask the model to write factual copy for you
number regions and hierarchy instead of hiding layout requirements in one paragraph
use exactly, only, and no other text to make the result auditable
change one core variable per test
keep the prompt, source data, output, and review notes together
VibeArt's infinite Canvas is useful for this broader workflow even before Qwen Image 3 is integrated: keep source material, model attempts, comparisons, edits, and approvals visible instead of losing them across separate chat threads.
Qwen Image 3 Vs GPT Image 2: A Practical Choice
The answer depends on the job, not a universal score.
Choose Qwen Image 3 for evaluation when you specifically need to test very long briefs, dense document-like composition, mixed scripts, and Qwen's interface or knowledge claims. Choose GPT Image 2 inside VibeArt when you need a model that is already available in the Canvas workflow for generation, comparison, refinement, and export.
Do not compare a cherry-picked official Qwen gallery with an average GPT output. Use the same source prompt, disable rewriting when the product allows it, repeat the test, record actual output dimensions and time, and score exactness separately from visual quality.
Yes. We generated all thirteen retained outputs in the official Qwen Chat web product with a signed-in account. Product access can vary by account and rollout state. A fresh image conversation defaulted to Qwen-Image 2.0 in our session, so verify and explicitly select the visible Qwen-Image 3.0 label before treating an output as a 3.0 result.
Is there a Qwen Image 3 API?
Yes, with an important qualification. Alibaba Cloud documents the model ID qwen-image-3.0-pro, but the model is currently an invite-only preview, not a generally available production API. The documented China-mainland rate limit is one request per minute. Treat both access and limits as preview conditions that may change.
Is Qwen Image 3 open source?
We could not verify released 3.0 weights or a 3.0 license. The older Qwen-Image repository is Apache-2.0, but that does not automatically apply to an unreleased 3.0 weight package.
How much does Qwen Image 3 cost?
Alibaba Cloud's pricing page lists qwen-image-3.0-pro as limited-time free in China mainland during the preview. That is not a stable general-availability price and does not prove that every region, account, or Qwen Chat tier is free. Check the current regional pricing page before building a cost model.
Does Qwen Image 3 generate 2K or 4K?
We could not verify a fixed model maximum. Record the actual dimensions from the product entry point used for each test; do not turn one product result into a universal model limit.
Final Verdict: Treat It Like A Generator Plus A QA System
Qwen-Image-3.0's most important idea is not prettier imagery. It is that one generated image can behave like a complex deliverable.
That idea is useful only when paired with a review system. Keep the source facts outside the bitmap. Audit exact copy and visual hierarchy independently. Record failed attempts. Preserve edit inputs. Measure repair time, not just generation time.
If Qwen Image 3 turns a 40-minute layout task into a five-minute generation plus a five-minute audit, it is a productivity tool. If a dense output takes longer to verify and repair than rebuilding the document in an editable format, it is a compelling demo rather than a deliverable.
That is the standard these thirteen retained Qwen-Image 3.0 first attempts measured.