- Most Midjourney content misses the point
- My workflow overview: Discord vs. web, organization, iteration
- Use case 1: Hero images for web projects
- Use case 2: Concept art for client pitches
- Use case 3: Social media visuals at scale
- Use case 4: Product mockups and UI inspiration
- The prompt formula that actually works
- What Midjourney still can't do reliably
- v6 vs v7: What actually changed in practice
- Is the Standard plan ($30/mo) worth it? My honest answer
- Tips for consistent results across a project
Most Midjourney content misses the point
If you search for Midjourney tutorials, you'll find two kinds of content. The first is jaw-dropping showcase galleries — dragons, fantasy landscapes, hyper-realistic portraits. Beautiful, sure. But unless you're selling desktop wallpapers, none of it translates to real work.
The second kind is prompt-engineering rabbit holes. Hour-long videos about the difference between --chaos 20 and --chaos 50, endless debates about whether cinematic lighting or dramatic lighting gets you closer to some aesthetic ideal.
Neither of these is the Midjourney I use daily. My version is a production tool. It produces images that go into live websites, client proposals, social media campaigns, and product documentation. It saves real hours. It has real limitations. And it requires a specific workflow that nobody seems to write about honestly.
This article is what I wish existed six months ago when I started treating Midjourney as a serious part of my creative process.
My workflow overview: Discord vs. web, organization, iteration
Midjourney runs in two places: the original Discord bot and the newer web interface at midjourney.com. If you're doing production work, use the web interface. Discord is chaotic, images get buried, and organizing anything is a nightmare. The web interface gives you a proper gallery, folders, and the ability to revisit and re-run prompts cleanly.
My organization system is simple: one folder per project. I name them after the client or deliverable — fintech-rebrand-2026, blog-hero-images-q2, social-campaign-spring. Within a project folder, I keep a plain text file (outside Midjourney) with the "anchor prompt" — the core style description I return to for every image in that set.
The iteration process matters more than the initial prompt. I rarely use a first-generation image directly. My typical flow:
- Generate 4 variations from the base prompt
- Pick the one closest to the direction I want, even if it's not quite right
- Use "Vary (Subtle)" to get 4 tighter variations around that starting point
- Upscale the winner
- If needed, use "Vary (Region)" to fix specific areas without regenerating the whole image
This flow means I'm rarely more than 3 rounds away from something usable. When I see people complaining that Midjourney "never gives me what I want," they're usually bailing after round one. The tool is iterative by design.
Use case 1: Hero images for web projects
Landing page hero images are where Midjourney has saved me the most time. Stock photos are either generic to the point of meaninglessness or expensive for anything distinctive. Custom photography requires a shoot. Midjourney is neither of those things.
The key challenge with web heroes is consistency across multiple images. A landing page might need a hero, three feature section images, and a team section background. They all need to feel like they belong to the same visual language. Without some mechanism for consistency, you'll generate five images that each look great individually but clash with each other on the page.
That mechanism is --sref (style reference). You generate one image you love, then use its URL as a style anchor for every subsequent image in the set:
a minimalist workspace with a laptop, soft morning light, plants in background
--sref https://cdn.midjourney.com/[your-anchor-image-url]
--ar 16:9 --v 6.1
The --sref parameter tells Midjourney to match the color palette, lighting mood, and overall aesthetic of the reference image, even when the subject changes. It's not pixel-perfect, but it's close enough that images generated with the same sref look like they came from the same photoshoot.
For a SaaS landing page rebrand, I generated 8 hero and section images in under an hour. All used the same --sref anchor with slightly different subjects. The client's designer said they looked "professionally shot." Total cost: about 40 image generations on the Standard plan.
For hero images specifically, I always use --ar 16:9 for desktop and generate a separate --ar 4:3 crop for mobile. The subjects and composition change enough between aspect ratios that a simple crop never works perfectly.
Use case 2: Concept art for client pitches
This is the use case that has paid for my Midjourney subscription many times over. Before committing to a designer — which costs real money and takes real time — I use Midjourney to put ideas on paper for client approval.
The pitch is simple: "Here are three visual directions we could take your brand. Which one resonates?" What used to require a designer spending a day building mood boards now takes me 30 minutes.
I generate 8-12 images for each direction, curate down to the 3-4 strongest, and present them in a simple PDF. Clients can react to something concrete. They stop saying "I'll know it when I see it" because now they can actually see something. The conversations become specific: "I like this color palette but the typography feels too corporate" instead of abstract color preference discussions that go nowhere.
The important caveat: I'm explicit with clients that these are directional, not final. Midjourney is showing us a flavor — a feeling, a palette, an energy. The actual deliverables will be crafted by a human designer. This distinction matters both ethically and practically. Nobody should be confused about what they're approving.
When a direction gets approved, I export my anchor images and prompts to pass to the designer as a brief. Instead of a vague style guide, they get concrete visual references and know immediately what aesthetic territory they're working in.
Use case 3: Social media visuals at scale
Social media has a voracious appetite for images. A consistent posting schedule across Instagram, LinkedIn, and TikTok can easily require 30+ unique visuals per month. At any realistic stock photo budget or designer rate, that math doesn't work for most small teams or solo operators.
Midjourney makes it work. The approach:
- Batch by aspect ratio. Do all your
--ar 9:16(Stories, TikTok) in one session, then all--ar 1:1(feed posts) in another. Switching between ratios mid-session breaks your mental flow and makes it hard to maintain consistency. - Use a brand sref throughout the month. Pick one "master" image that captures your brand energy. Use it as your
--srefanchor for every social image that month. Your feed starts looking cohesive without anyone manually forcing it. - Batch thematically. If you're doing a product launch campaign, generate all the launch images in one sitting. Same session, same sref, same lighting notes in the prompt. Doing this across multiple sessions introduces visual drift.
One workflow I've found particularly effective: generate 16 images in a session (4 prompts x 4 variations), curate down to the 8 best, and schedule them two weeks out. By the time those run out, you do another batch. The creative overhead is one focused hour per two weeks rather than constant daily scrambling.
--ar 9:16 — Instagram/TikTok Stories, Reels vertical | --ar 1:1 — Instagram feed, LinkedIn posts | --ar 16:9 — Twitter/X header, YouTube thumbnail, LinkedIn articles | --ar 4:5 — Instagram portrait posts (performs well in feed)
Use case 4: Product mockups and UI inspiration
This one comes with the clearest disclaimer: Midjourney is not a UI design tool. It cannot produce a pixel-perfect mockup you hand off to developers. It can't generate real interface components. It definitely can't output anything with accurate text labels.
But it's excellent for something more important in the early stages: direction-setting.
When I'm exploring what a product interface could feel like, Midjourney can generate 20 visual impressions of "a minimal dark-mode dashboard with warm amber accents" in the time it would take me to open Figma. I'm not extracting pixels from these images. I'm answering the question "does this aesthetic deserve to be explored further?" before investing the hours to actually build it.
The same logic applies to physical product mockups. If you're evaluating packaging options for a product, Midjourney can show you what matte black minimalist packaging versus a colorful illustrated approach would look like at a general level. The printing specifications come later. The directional decision — which often saves you from going down an expensive wrong path — can happen in an afternoon.
I also use Midjourney for UI inspiration rather than specification. If I'm stuck on how to present a data-heavy feature, generating a few impressionistic "data visualization dashboard with [style X]" images can break the creative block. I'm not copying anything; I'm unsticking my thinking.
The prompt formula that actually works
After thousands of generations, I've converged on a consistent prompt structure. Here it is:
[subject] [style reference] [mood/lighting] [--ar X:X] [--v 6.1]
Let me break that down with real examples:
# Hero image for a fintech landing page
a person reviewing financial charts on a laptop,
clean editorial photography style, warm morning light,
shallow depth of field, neutral background
--ar 16:9 --v 6.1
# Social post for a productivity app launch
flat lay of notebook, coffee, and phone showing a minimal app interface,
modern lifestyle photography, soft natural light, sage and cream color palette
--ar 1:1 --v 6.1
# Concept art for a meditation app
abstract flowing shapes representing calm and stillness,
watercolor illustration style, muted blues and greens, serene atmosphere
--ar 4:5 --v 6.1 --stylize 200
A few notes on what I've learned:
- Lead with the subject. Midjourney weighs earlier words more heavily. The first 5-10 words have outsized influence.
- Name a photography or art style explicitly. "Editorial photography," "lifestyle photography," "flat lay," "conceptual illustration" — these get you much more targeted results than describing visual properties abstractly.
- Lighting is the fastest way to change mood. "Warm morning light" vs. "dramatic side lighting" vs. "soft overcast" produce fundamentally different feels even with identical subjects.
- Always specify
--v 6.1(or--v 7) explicitly. The default version can vary, and you want consistent results across a project. - Keep prompts under 60 words for most work. Longer prompts often introduce conflicts that the model resolves unpredictably.
What Midjourney still can't do reliably
Honesty matters here. After six months of production use, these are the failure modes I've learned to work around:
- Text in images. Any readable text in a Midjourney image is a gamble. It looks like text, often in approximately the right place, but the actual letters are frequently mangled. Never rely on Midjourney for images where readable text matters. Add text in post using Figma, Canva, or Photoshop.
- Hands with more than four fingers. The infamous problem. v6 is significantly better than v5, and v7 better still, but complex hand positions remain unreliable. If hands are central to your image, plan to use "Vary (Region)" to fix them or regenerate until you get lucky.
- Specific logos and brand marks. Midjourney will approximate a recognizable brand aesthetic, but it cannot reliably reproduce a specific logo. Don't ask it to. Composite brand elements in post if needed.
- Exact facial consistency. If you need the same person to appear across 10 images, Midjourney's character reference feature (
--cref) helps but isn't reliable enough for anything requiring true likeness consistency. Use a real model or character illustration tool instead. - Technical diagrams or information design. Charts, flowcharts, architectural diagrams — Midjourney produces things that look like these but contain no real information. The data is fabricated. Never use Midjourney for anything that needs to convey accurate technical information.
v6 vs v7: What actually changed in practice
v7 launched in early 2025 and the marketing around it was breathless. Having used both extensively for production work, here's my honest assessment:
Where v7 is genuinely better: Photorealism. If you're generating images that need to look like photographs — product shots, lifestyle images, environmental photos — v7 is noticeably more convincing. The lighting physics are more accurate, textures are more detailed, and that uncanny-valley AI sheen is significantly reduced.
v7 also handles hands and fingers better. Not perfectly, but the failure rate on natural hand positions dropped substantially. For images where hands are visible, v7 is worth defaulting to.
Where v6.1 still wins: Stylized and illustrative work. The more you push toward painterly, graphic, or illustrative aesthetics, the less the photorealism improvements of v7 matter — and v6.1's output for stylized work often feels more intentional and less like a photograph that's been artistically filtered. For concept art, illustration-style social images, and anything with a strong graphic aesthetic, I still often reach for v6.1.
Speed and cost: v7 uses more GPU time per generation, which eats through your fast-generation hours faster. If you're doing high-volume batch work and photorealism isn't a priority, the efficiency argument for v6.1 remains real.
My default approach: use v7 for photography-adjacent work and v6.1 for everything illustrative. A few minutes testing both on a new project brief is always worth it before committing to one for a full session.
Is the Standard plan ($30/mo) worth it? My honest answer after 6 months
The short answer: yes, but only if you actually use it.
The Standard plan gives you 15 hours of fast GPU time per month. In practice, that's roughly 900-1,200 image generations depending on which model and settings you're using. If you're doing real production work — multiple client projects, ongoing social media needs, regular web content — you'll use most of that. If you're experimenting occasionally, the Basic plan ($10/mo) is probably enough.
The key advantage of Standard over Basic isn't just more fast hours. It's unlimited relaxed generations. Relaxed mode is slower (can take a few minutes per job vs. seconds for fast), but the images are identical quality. My workflow: I use fast mode for active iteration when I'm in a flow and need quick feedback. For batch generation of a full set where I'm not watching the screen, I switch to relaxed and let it run while I work on something else.
The Pro plan ($60/mo) adds stealth mode (private generations) and more fast hours. For client work with confidential briefs, stealth mode is worth paying for. If your work is public-facing and you don't need privacy, Standard is the ceiling you need.
At the start of a new project, do your exploratory generation in fast mode. Once you've nailed the anchor prompt and style reference, switch to relaxed for the bulk generation. You'll conserve fast hours for when iteration speed actually matters.
Compared to alternatives: DALL-E 3 (included with ChatGPT Plus at $20/mo) is more accessible but produces less photorealistic results and has less precise style control. Ideogram is notably better at text in images but lags behind Midjourney for photographic and stylized work. If text-in-image is your primary use case, Ideogram deserves a serious look. For everything else, Midjourney remains the benchmark.
Tips for getting consistent results across a project
Consistency is the hardest problem in production Midjourney use. Here's the toolkit I've built up:
- Save your anchor prompt. Create a text file for each project with the exact base prompt, sref URL, and parameter flags. Copy-paste this at the start of every session. Never reconstruct it from memory.
- Use
--srefreligiously. Generate one image you love in the first session. Save its URL. Use it as the style reference for every subsequent image in the project. This single practice has more impact on visual consistency than any amount of prompt tuning. - Fix your seed for exploratory batches. Use
--seed [number]to reproduce a starting point when you want to explore variations around a specific composition. Note: seeds are version-specific, so changing from v6 to v7 will produce different output even with the same seed. - Do all images for a deliverable in one session. Your prompts will naturally evolve slightly across sessions. If you generate half a set on Monday and the other half on Friday, they'll look different in subtle ways. Same-session batching is the simplest consistency guarantee.
- Accept that perfect consistency is impossible, and design around it. Even with identical prompts and the same sref, Midjourney generates stochastically. The right design approach is to select images with similar energy, not images that are pixel-identical. Your layout and typography will do more for visual cohesion than any amount of prompt engineering.
After six months, my honest relationship with Midjourney is this: it's a tool that makes me significantly faster at the visual parts of my work, within a specific lane. It doesn't replace creative judgment. It doesn't replace craft. It doesn't work for everything. But within that lane — hero images, concept exploration, social batches, directional mockups — it's become genuinely indispensable. The key is treating it as a production tool with a workflow, not a toy you play with until something impressive appears.
Read the full Midjourney review for a complete breakdown of pricing, features, and how it compares to other image AI tools.