
Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget




Welcome to the official repository for Boogu-Image-0.1 !
English | 中文
⚠️ Important Notice
The Boogu team does NOT currently provide any paid API, subscription, or commercial service for Boogu-Image. Any paid product or service offered under the name "Boogu-Image" — or any similar / variant name such as booguimage, Boogu Image, Boogu, etc. — is NOT affiliated with this project and is unofficial. Please verify carefully before making any payment, and stay vigilant to protect your personal privacy and financial safety.
Boogu-Image-0.1 is a research project only, and not an official model release.
📖 Introduction
Boogu-Image-0.1 is a competitive Apache-2.0 open-source unified image generation and editing model family, including Base, Turbo, Edit, and Edit-Turbo, and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing, and Chinese-English text rendering. Closed-source multimodal understanding and generation systems like Nano Banana Pro and GPT-Image-2 achieve remarkable performance not because of a single model, but through a highly unified suite of system capabilities. However, under training compute that is extremely limited compared with closed-source systems, we find that systematically improving a model's understanding ability, data quality, and training pipeline can still significantly improve image generation and editing performance. Specifically, compared with some existing open-source models, our training data scale is roughly one order of magnitude smaller. We hope our empirical study and open-source release will help advance the open-source ecosystem for multimodal generation and understanding.
This repository provides checkpoints and inference code for Boogu-Image-0.1.
📣 News
- 2026-07-23 🔥 NPU support is now available! Check out the
npu branch for initial NPU backend support and instructions. We welcome feedback and bug reports!
- 2026-07-22 🔥 vLLM-Omni now supports Boogu-Image! You can now run Boogu-Image models with the vLLM-Omni inference server. See the official recipe for details and setup: vllm-project/vllm-omni Boogu-Image recipe.
- 2026-07-16 🔥 🚀 🚀 The Boogu-Image-0.1 Technical Report is released! We have released our technical report. If you find our series of models or the report helpful to your research or applications, please consider citing our paper. Thank you very much for your support! You can read the paper here: arXiv:2607.13125.
- 2026-07-08 🔥 Boogu-Image-0.1-Edit-Turbo (Image-to-Image hotfix) is released! This update addresses several issues in the previous version, including severe image quality degradation and poor performance on removal and other editing tasks. The new checkpoints are released on Hugging Face in the revisions hotfix-1k-20260708 and hotfix-1k5-20260708. We recommend downloading the 1K checkpoint for more stable results. Try the online demo for 1K res and online demo for 1.5K res.
- 2026-06-30 🔥 Boogu-Image-0.1-Edit-Turbo (Image-to-Image) is released! Four-step distilled variant of the editing model for fast inference. Try the online demo for 1k res and online demo for 1.5k res.
- 2026-06-25 🔥 Boogu-Image-0.1-Turbo-hotfix (Text-to-Image) is now online! The new checkpoint is released on Huggingface in the revision hotfix-20260625. This is a minor patch release. We fixed visual artifacts appear in different aspect ratio, background overfitting artifacts, and other artifacts. Model weights are updated, no feature changes.
- 2026-06-17 🔥 ComfyUI-Boogu powered by ComfyUI is released! Thank you, ComfyUI!
- 2026-06-17 🔥 ComfyUI-Boogu is released!
- 2026-06-16 🔥 Boogu-Image-0.1-Base (Text-to-Image) is released! The core text-to-image foundation model. Try the online demo.
- 2026-06-16 🎨 Boogu-Image-0.1-Edit (Image-to-Image) is released! Image editing and transformation capabilities now available. The model supports resolutions up to 2K, but the results are more stable at 1K. Try the online demo. Only support 1 reference image for now. Will try our best to support more reference images. Stay tuned! Boogu-Image-0.1-Edit on single-image editing is strong. More failure cases are welcome.
- 2026-06-16 🚀 Boogu-Image-0.1-Turbo is released! Four-step distilled variant for fast inference and photorealistic generation. Try the online demo.
🌐 Join the Community!
Boogu-Image is built to grow with its users. Share what you create, report issues, exchange ideas, and help shape what comes next.
WeChat Group

🏆 Boogu Arena
Since we could not evaluate on LM Arena directly, we built Boogu Arena, an LM Arena-style preference evaluation. We use an LLM to generate diverse user personas, then ask each persona to produce image generation prompts, resulting in 1K+ test prompts that we will release publicly for community reproduction. The ELO leaderboard below spans leading closed- and open-source systems. We welcome teams with questions about the results to contact us so that we can work toward a more objective, fair, and reproducible evaluation.

✨ Highlights
- 📸 Beautiful and Precise Photography — Accurately understands photography prompts and generates high-quality images with natural lighting, coherent composition, and faithful details, preserving coherent subject, background, and spatial relationships even in complex real-world scenes

- 📝 Diverse and Stable Text Rendering — Supports a wide range of text-heavy designs — posters, stamps, documents, interfaces, brand guides, and handwritten boards — with readable structure, stable typography, and robust bilingual (Chinese/English) rendering across diverse layouts

- 🎨 Diverse and Beautiful Stylization — Handles stylized generation across miniature 3D scenes, Chinese-inspired gilded aesthetics, shining fantasy visuals, anime portraits, and mythic character art — not just style transfer, but stable, attractive, and prompt-aware creative generation

- 🖌️ Versatile Image Editing — Handles a wide spectrum of editing tasks, including object insertion, replacement and removal, attribute and material modification, background and scene replacement, and faithful style transfer across artistic looks, while keeping the source subject and composition coherent

- 🪧 Personalized Poster Design & Product Rendering — Generates personalized poster layouts and clean product visualizations with consistent branding, refined typography, and product-grade lighting and composition

- ✍️ Precise Text Editing — Enables fine-grained, in-image text editing — replacing, adding, or removing characters in both Chinese and English — and flexibly adapts fonts, weights, colors, and layouts to match different design intents

- 📊 Competitive General Performance — Demonstrates competitive performance across many scenarios and benchmarks, with the Boogu-Image-0.1 family ranking among the very top of evaluated open- and closed-source systems in Boogu Arena
📖 For the full set of practical lessons and an honest account of current limitations, see Responsible AI & Limitations below.
🔬 Scenario-wise Comparison
Beyond overall arena rankings, we break performance down by scenario across leading open-source peers. Ratings reflect our internal evaluation of typical prompts in each category.
| Model |
Realistic Photography |
Simple Text Rendering |
Dense Text Rendering |
| Boogu-Image-0.1-Turbo |
⭐⭐⭐⭐ |
⭐⭐⭐⭐ |
⭐⭐⭐ |
| Boogu-Image-0.1-Base |
⭐⭐⭐ |
⭐⭐⭐⭐ |
⭐⭐⭐⭐ |
| Z-Image-Turbo |
⭐⭐⭐⭐ |
⭐⭐⭐ |
⭐⭐ |
| Qwen-Image-2512 |
⭐⭐⭐ |
⭐⭐⭐⭐ |
⭐⭐⭐ |
- 📸 Photography with reliable text rendering — Boogu-Image-0.1-Turbo delivers realistic photography, while also offering solid performance on both simple and dense text rendering.
- 📝 Strong dense text rendering — Boogu-Image-0.1-Base shows competitive results on dense, layout-heavy text scenarios such as posters, documents, brand guides, and complex bilingual designs.
- 💡 Recommendation — When your workload is dominated by dense / ultra-dense text rendering needs, we recommend running Boogu-Image-0.1-Base at 2K output resolution for the best layout fidelity and character accuracy.
📥 Model Zoo
| Model |
Params |
Training |
Steps |
CFG |
Task |
Hugging Face |
ModelScope |
Demo |
| Boogu-Image-0.1-Base |
10B |
Joint Training |
25~50 |
2.0~5.0 |
|
|
|
|
(e.g., 4.0) | T2I |
|
|
|
| Boogu-Image-0.1-Base-fp8 | 10B | Joint Training | 25~50 | 2.0~5.0
(e.g., 4.0) | T2I |
|
| — |
| Boogu-Image-0.1-Edit | 10B | Joint Training | 25~50 | 2.0~5.0
(e.g., 5.0) | TI2I |
|
|
|
| Boogu-Image-0.1-Edit-fp8 | 10B | Joint Training | 25~50 | 2.0~5.0
(e.g., 5.0) | TI2I |
|
| — |
| Boogu-Image-0.1-Turbo | 10B | + Decoupled DMD | 4 | 1.0 | T2I |
|
|
|
| Boogu-Image-0.1-Turbo-fp8 |