How I Pick the Right AI Model for Each Job
For a long time I had one rule for AI: whichever tool I'd opened most recently was the one I used for everything. If ChatGPT was the last tab, I wrote landing copy in ChatGPT. If Claude had been open since morning, I debugged code in Claude. The output was always *fine*. None of it was great.
The shift came when I started running the same prompt across three or four models and comparing the results. Not in some rigorous benchmark way — just side by side, on real work I had to ship. The pattern was unmissable. Every model has a shape. Use it for the job that fits its shape and the output is dramatically better. Use it for the job that doesn't and you spend an hour editing what should have taken ten minutes.
I now have a routing rule for every kind of work I do. It looks roughly like this.
What I use Claude for
Long-form writing. Editorial judgment. Anything where the output has to read well, sustain an argument over more than two paragraphs, and avoid the AI tell.
Claude is the only model I trust to write a 1,000-word blog draft I can actually edit into something publishable. Other models can produce 1,000 words on demand. The difference is what the 1,000 words feel like. Claude has the rhythm closer to how a human writes — uneven paragraph length, willing to hedge, less compulsive about parallel constructions and tidy summaries.
I also use Claude for any prompt where I'm asking the model to push back on me. Decisions, design reviews, copy critique. It's better at being direct without being performatively contrarian.
This is the prompt I use for most editorial work:
I'll paste a draft. You're an editor with strong taste who's read enough internet writing to know what sounds good and what sounds generic. Don't rewrite it. Tell me: (1) the one paragraph that should be cut, (2) the one sentence that's load-bearing and shouldn't change, (3) one specific edit that would make the opening sharper. Be direct. Don't soften.
Category: writing. Tool: Claude. Status: favorite.
What Claude is *not* great at: image generation prompts where you need exact technical jargon, fast iteration on code where you want to bang out twenty variations in five minutes, anything in a chat workflow where speed of response matters more than depth of response.
What I use ChatGPT for
Code work. Quick iteration. Anywhere the value is in volume of attempts rather than quality of any single attempt.
ChatGPT is faster to think out loud with, faster to dump messy context into, and the code suggestions tend to land more often than not for the kind of work I do (Next.js, TypeScript, Supabase). When I'm debugging at 11pm and I want to paste an error message and a stack trace and get five things to try, ChatGPT is the tab I open.
I also use it for anything that needs to feel snappy. Quick summaries. Rewording an email I'm about to send. Two-line social posts. The output isn't always better than Claude — but the *speed* of the round trip matters when you're using AI as a thinking partner rather than as a writer.
This is my go-to debugging prompt:
I'm hitting this error in [STACK]: [ERROR + STACK TRACE]. Here's the relevant code: [PASTE]. Don't tell me what the error means — I can read. Give me the three most likely causes ranked by probability, and the one-line check I should run for each.
Category: code. Tool: ChatGPT. Status: favorite.
What ChatGPT is *not* great at: long-form anything. Past about 600 words the writing flattens out into the AI rhythm and you can't edit it back into something good. It's also worse at saying "I don't know" — Claude will hedge, ChatGPT will confidently produce a plausible answer.
What I use Midjourney for
Anything where the image needs to look composed. Product shots, blog headers, mood boards, anything that has to live next to designed work without looking obviously generated.
Midjourney's defaults are closer to what a designer would call "good" than any other image model I've used. Lighting, composition, depth of field, color palette — the model has opinions, and the opinions are usually right. The cost of those opinions is that it's harder to make Midjourney do something *specific*. If I want a photo of my product on a particular surface in a particular light, I can dial it in. If I want a precise illustration with text and arrows, Midjourney will fight me.
This is a Midjourney prompt I reuse for product photography:
Cinematic product photography. [SUBJECT] on a [SURFACE], [LIGHTING], shallow depth of field, soft shadows, [COLOR PALETTE]. Shot on a 50mm lens. No text, no logos, no people. High detail, photographic realism, 4k.
Category: image. Tool: Midjourney. Status: favorite.
What Midjourney is *not* great at: anything with text, anything where consistency across multiple images matters (a recurring character, a brand mascot), anything where the user wants strict compositional control. For those I switch to other image tools that trade taste for steerability.
What I use Sora for
Video where I want something that *feels* cinematic without setting up a shoot. B-roll, atmospheric clips, mood-setting shots that intercut with real footage.
Sora has the same kind of inherent taste as Midjourney — its defaults look like film, not like AI. The output isn't always usable as standalone video, but as a clip that lives inside a larger edit, it's transformed how fast I can produce a video. A B-roll shot I would have spent an afternoon hunting for in stock footage I can now generate in eight minutes.
This is my Sora prompt for atmospheric B-roll:
8-second handheld shot. [SUBJECT] in [ENVIRONMENT], [TIME OF DAY], [WEATHER]. Slow camera drift to the right. Shallow depth of field. Slight film grain. No text, no on-screen graphics, no people in focus. Cinematic color grade, muted tones.
Category: video. Tool: Sora. Status: tested.
What Sora is *not* great at: dialogue, faces in close-up, anything longer than about ten seconds where motion has to stay coherent. I treat it as a generator of cuts, not full scenes.
What I use Suno for
Music beds. Demo tracks. Anything where I need to put audio under a video and royalty-free libraries don't have the exact mood.
Suno is the model that surprised me the most. The first time I used it I assumed I'd get something obviously synthetic, suitable for a tech demo and nothing else. The output was good enough that I now use Suno tracks under client work without flagging it. Genre control is precise. Mood control is precise. Mix quality is good enough that you can lay it under voiceover without it fighting the dialogue.
This is the prompt structure I use:
[GENRE], [MOOD], [TEMPO]. Instrumentation: [INSTRUMENTS]. Structure: intro, build, drop, outro. Length: ~2 minutes. Production: [MIXING NOTES]. No vocals.
Category: music. Tool: Suno. Status: favorite.
What Suno is *not* great at: tracks with vocals (the lyrics are still rough), anything where you need exact timing to picture, anything you'd actually release as a standalone song rather than as a bed under other content.
The cost of using the wrong tool
The reason this matters: using the wrong model isn't free. The output is not just "slightly worse." It's worse in a way that costs you editing time, that ships subtly off-brand, that you don't notice until you compare it to what you would have gotten from the right tool.
I had a stretch where I was writing all my landing copy in ChatGPT because that was my open tab. Every page had a faint sameness. I assumed the problem was me. The day I rewrote three pages in Claude using the same outlines, the difference was so obvious it felt embarrassing. The pages weren't dramatically different in *content*. They were different in *rhythm* — and rhythm is most of what makes copy convert.
Similarly, I spent weeks trying to get a more steerable image model to produce a hero image for the Super Prompts homepage. Every output was technically correct and aesthetically dead. I switched to Midjourney, used a one-line prompt, and had something usable in three generations.
Picking the right model is not a small optimization. It's the difference between AI as a tool that compounds your output and AI as a treadmill that produces mediocre work faster.
How to figure out your routing
If you have time for one experiment, do this: pick a piece of work you do regularly — landing copy, debugging, image generation, whatever. Run the same prompt on three different models. Don't try to rate them in the abstract. Just look at the outputs side by side and ask which one you'd ship as-is, which one you'd edit, and which one you'd throw out.
Do that three or four times across different jobs and your routing rule writes itself. You'll find the model that fits your taste for each kind of work, and the cost of using the wrong one will become obvious enough that you stop doing it.
The model isn't the magic. The match between the model and the job is.