Stable Diffusion

  • Stable Diffusion (Stability AI, London 2018 model 2022, Active, Not Shut Down)

    Status: Active (Stability AI founded 2018 London by Emad Mostaque, SD 1.0 released Aug 2022, latest open weights SD 3.5 family​ + SDXL / SDXL Turbo​ 2024–2025; weights on Hugging Face, code on GitHub, DreamStudio/API live 2026-08, no discontinuation). Category ai-productivity-data; pipeline_stage=none. Not a SaaS product, not a chat image button​ — it is the open-weight latent-diffusion image foundation model you self-host or call via API, with the largest modding ecosystem (LoRA / ControlNet / ComfyUI / A1111) in the lane.

    1.  Introduction

    Stable Diffusion is the Linux of text-to-image: a permissively-licensed latent diffusion model that turns prompts (and images) into pictures entirely on your own GPU or inside your own backend. Core lineage runs SD 1.5 → SD 2.1 → SDXL 1.0 (1024²) → SD 3.5 Large/Turbo/Medium (MMDiT, ~1 MP), with sibling models Stable Image Ultra/Core, SDXL Turbo (1-step distill), Stable Video / Fast 3D / Stable Audio in the wider family. You drive it through Automatic1111 WebUI, ComfyUI​ (node graphs), Fooocus​ (Midjourney-like simplicity), or the Stability API / DreamStudio. Because weights are open, you can fine-tune on your brand, lock pose with ControlNet, inject style with LoRA, inpaint/outpaint locally, and ship unlimited images at zero per-call cost. It trades out-of-the-box polish (Midjourney) and text-in-image accuracy (DALL·E 3 / Flux) for total ownership.

    2. Core Features

    Layer Capability Notes
    Generation modes Text-to-image, Image-to-image, Inpainting, Outpainting, Upscale, Sketch/Structure/Style conditioning SDXL 1024², SD3.5 up to 1 MP
    Local / private Run on NVIDIA 8 GB+ VRAM (12 GB rec), Apple Silicon, select AMD; images never leave machine RTX 3060 12 GB ~$200 handles most
    Ecosystem 90,000+ community checkpoints on Hugging Face/Civitai; LoRA, Embeddings, ControlNet, AnimateDiff, ADetailer The actual moat
    UIs A1111 (params), ComfyUI (nodes/reproducible), Fooocus (easy), InvokeAI, Forge No official “app”, assemble your own
    API / cloud Stability AI API (Stable Image Ultra/Core, SD3.5, SDXL), DreamStudio credits, Bedrock/NVIDIA endpoints, Replicate/RunPod 1 credit ≈ $0.01
    Modality spillover Stable Video Diffusion, Stable Fast 3D, Stable Audio 2.5/3.0, Virtual Camera Same brand, separate weights
    No No built-in content moderation, no chat wrapper, no mobile app, no turnkey commercial indemnity Pair with guardrails for prod

    3. Pricing (2026-08)

    Path Cost Gate
    Self-host (Community License) $0​ model + GPU electricity Free for individuals / orgs < $1M annual revenue; unlimited generations; GPU capex $200–1600 one-time
    DreamStudio $10 / 1,000 credits​ (~5,000 basic images) No GPU, browser, trial credits on signup
    Stability API Pay-as-you-go ~$0.002–0.06/image​ (SDXL 0.9 cr, SD3.5 Large 6.5 cr, Ultra 8 cr, 1 cr=$0.01) Brand Studio Creator $19/mo 2k credits, Core $50/mo 5k credits
    Cloud GPU rental RunPod/Vast.ai $0.20–0.80/hr No local GPU
    Enterprise License Custom (orgs ≥ $1M revenue​ must sign) SLA, consulting, custom training, indemnity

     

    pricing_type=open-weight-free + credit-api + gpu-capex; starting_price=$0; the model is free, the bill is your GPU or your API calls.

    4. Access Type

    🌐 Hugging Face weights · GitHub code · A1111/ComfyUI/Fooocus local · DreamStudio web · Stability REST API + SDKs · Bedrock/Replicate/RunPod · no official mobile app, no chat UI.

    5. Reviews (2026)

    • ShipSquad / LaunchToolsAI / SaaS Radar: 4.3–4.6/5 — “unbeatable control & cost at scale, steep setup, default aesthetics trail MJ”.
    • AIMagicX blind test (600 imgs): SD3.5 product/person consistency 55%, ~5 s/gen, $0.01/api — cheaper but more rerolls than Flux/MJ.
    • G2/Capterra: sparse (it’s a model, not a SaaS) — sentiment mirrors editorial.
    • Verdict: right answer for devs/studios shipping image gen into a product or generating 1k+/mo privately; wrong answer for “type a sentence, get a poster”. Gripes = GPU barrier, sampler/CFG/LoRA soup, SD3 license stricter than SD1.5/SDXL, hands/text still weaker than Flux/DALL·E 3, no safety net.

    6. Best For / Not For

    Best For Not For
    Devs embedding image gen in apps (API or self-host, no per-seat tax) Non-technical users wanting Midjourney-grade output in 1 click
    Studios generating 1k–100k/mo privately (product, training data, variants) Teams needing built-in IP indemnity + moderation (→ Adobe Firefly)
    Fine-tuners / character-consistency / brand-locked style (LoRA+ControlNet) Perfect text-in-image / menu / infographic (→ DALL·E 3 / Ideogram / Flux)
    Privacy-bound orgs (health, defense, internal assets) No-GPU, no-CLI, no-Python tolerance
    Researchers experimenting with diffusion internals Casual social-media thumbnail once a week (overkill)

    7. Competitors (Lane Map)

    Lane Tool SD edge / gap
    Closed art king Midjourney V8 MJ prettier out-of-box, no local, $10–120/mo, no API
    Chat-native + text DALL·E 3 / GPT Image 2 Better prompt/text rendering, closed, $0.04–0.12/img
    Open-weight photoreal Flux 1.1 Pro / Dev / Schnell (Black Forest) Flux better faces/hands/text, similar open ethos, Dev weighs more VRAM
    Design-suite safe Adobe Firefly Firefly indemnified + PS tie-in, less pipeline control
    App-builders Leonardo.AI, Recraft v3, Ideogram 2 Nicer UI/styles/credits, closed weights
    Video sibling Runway, Kling, Luma Different modality; SD Video still early
    Self-host alt Kohya (SD fine-tune stack), ForgeUI Tooling, not rival model

     

Do Not Sell or Share My Personal Information Cookie Settings