# How to Vibe Code Your Own Midjourney (and Stop Paying for It)

> An independent research lab exploring new mediums of thought

- Site: https://midjourney.com
- Category: AI Media Generation
- Verdict: **Solid side project** (62/100 vibecodeable)
- Estimated effort: 2-3 weeks part-time

## Verdict

Build a personal prompt-generation dashboard wrapper using fal.ai for inference, but keep paying for Midjourney if you want its proprietary aesthetic model weights.

Replicating Midjourney's proprietary fine-tuned diffusion models and custom aesthetic weights from scratch is impossible for a solo developer. However, you can build a highly functional personal clone in a few weeks by combining a Next.js web frontend, a serverless database to track generations, and an inference API like fal.ai running open models like Flux. The real engineering friction will be managing asynchronous job queues, storing high-resolution image assets cheaply, and matching the seamless multi-action grid upscaling UX.

### What you can't replicate

- Midjourney's proprietary v6/v7 model weights and distinct aesthetic rendering
- Massive H100/A100 proprietary GPU infrastructure
- The 21-million-member Discord community and prompt dataset

## What it does

AI-powered text-to-image and video generation platform accessible via web dashboard and Discord bot.

### Core features

- Natural language prompt parsing and weighting
- Latent diffusion image generation pipeline
- Grid generation, upscaling, and variation actions
- Asynchronous job queue with priority routing
- Web application dashboard for gallery and prompt history
- User authentication and credit/GPU tracking
- Discord bot integration for command-based generation

## The business

### Pricing

- Basic: $10/mo — ~3.3 hours of Fast GPU time per month (~200 images)
- Standard: $30/mo — ~15 hours of Fast GPU time, unlimited Relax mode
- Pro: $60/mo — ~30 hours Fast GPU time, Stealth mode
- Mega: $120/mo — ~60 hours Fast GPU time

### Funding

$0 raised.

Founded 2021.
Team size: 40-160.

## The hard parts

- Managing high-throughput GPU cluster inference queues without crashing under burst traffic
- Training and hosting custom latent diffusion checkpoints that match proprietary aesthetic tuning
- Handling real-time state synchronization for long-running asynchronous jobs across web clients
- Optimizing VRAM utilization for simultaneous grid generations and high-res upscales

## How to vibe code Midjourney

### Prerequisites

- Node.js (free): Runtime environment for the Next.js full-stack framework.
- GitHub (free): Code repository hosting and deployment integration.
- fal.ai API Account (Usage-based (~$0.003/image)): Provides rapid image generation inference using Flux models.

### Recommended AI tools

- Claude Code: Autonomous terminal coding agent to scaffold the full-stack architecture and write prompt parsing logic.
- Cursor: AI-powered editor for polishing the dashboard gallery UI and grid management layouts.

### Stack

- Frontend: Next.js with Tailwind CSS and shadcn/ui
- Backend: Next.js API Routes / Server Actions
- Database: Turso (SQLite at the edge)
- Auth: better-auth
- Payments: None (Personal use clone)
- Other: fal.ai API (Flux inference), Cloudflare R2 (Image storage)

### Hosting

- Vercel (Hosting the Next.js web application frontend and API serverless routes.): $0/mo (Hobby tier)
- Cloudflare (R2 object storage for generated image assets with zero egress fees.): $0/mo (Free tier)

### Build guide

1. **Project Scaffolding and Database Schema** — Initialize a Next.js project with Tailwind CSS, configure better-auth for single-user authentication, and set up Turso SQLite database tables for prompts, generations, and user credits.

```
Scaffold a new Next.js 16 application with TypeScript, Tailwind CSS, and App Router. Set up better-auth with email/password authentication connected to a Turso SQLite database using Drizzle ORM. Create database tables for users, generations (id, prompt, negative_prompt, status, image_url, grid_layout, created_at), and user_credits. Include standard error boundaries and a clean folder structure.
```

2. **Prompt Engineering & Parameter Parser** — Build a parser utility that handles Midjourney-style parameters like aspect ratios (--ar 16:9), stylize (--s 250), and chaos (--c 10) to map them into API parameters.

```
Create a TypeScript utility function that parses Midjourney-style text prompts. It should extract parameters such as aspect ratio (--ar 1:1, 16:9, 4:3), stylize values (--s), chaos (--c), and version flags (--v). The function must return a clean prompt string for the AI model and a structured options object representing the generation parameters.
```

3. **Inference API Integration with fal.ai** — Implement the backend job initiation service that submits prompts and parameters to fal.ai's Flux inference endpoint and handles asynchronous webhook results.

```
Implement an asynchronous generation service in Next.js server actions that integrates with the fal.ai API using the Flux model. When a user submits a prompt, create a generation record with 'pending' status in the Turso database, call the fal.ai endpoint with parsed parameters, and handle polling or webhook completion to update the image URL and status.
```

4. **Dashboard UI & Grid Generation View** — Build the main prompt input bar, real-time generation grid view with loading skeleton states, and historical feed layout inspired by the Midjourney web app.

```
Create a responsive dashboard UI using Tailwind CSS and shadcn/ui components. Include a prominent prompt input bar at the bottom with parameter adjustment controls, a live-updating grid view for pending generations with loading skeleton states, and a main gallery feed displaying past generated image grids with filtering and search capabilities.
```

5. **Image Upscaling and Variation Actions** — Add UI buttons and backend handlers for individual image upscaling (U1, U2, U3, U4) and creating variations (V1, V2, V3, V4) from a 2x2 generation grid.

```
Build image manipulation UI components that overlay on completed 2x2 generation grids, providing Upscale (U1-U4) and Variation (V1-V4) action buttons. Implement backend endpoints that crop or re-prompt fal.ai using the parent generation context and selected grid quadrant index, storing the resulting single asset in Cloudflare R2 storage.
```

6. **Storage Persistence & Polish** — Configure Cloudflare R2 storage integration to download generated images from temporary inference URLs and store them permanently, adding final UI toast notifications and polish.

```
Configure an S3-compatible utility using `@aws-sdk/client-s3` to download completed images from fal.ai temporary URLs and upload them securely to Cloudflare R2 object storage. Update the database record with the permanent R2 public URL. Add toast notifications for job completion and error states across the dashboard.
```

### Cost vs paying

**Starting costs (one-time):**

- AI Coding Assistant Pro Subscription: $20
- fal.ai Initial Generation Credits: $10
- Total: ~$30 one-time

**Ongoing costs (monthly):**

- fal.ai API Usage (~200 images/mo): ~$3/mo
- Vercel & Cloudflare Hosting: $0/mo
- Total: ~$3/mo

- Paying for the SaaS instead: $30/mo (Standard Plan)
- Build time: 18 hours
- AI tool credits: $20 (Claude Code / Cursor Pro)
- Break-even: 1 month

## Sources

- [Midjourney System Design Guide](https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGpYO3Q4eFVNBo6TZjZzcPWFGWWqbdO9AJOf5uN7kBWtVJEKysj4HmbTH9tvIKbOtrpp2uTBvoVSa1qLYFckCoJCfUsyT1watSVSess1SXjG5DUV8CTNW0IQHKmguDJbtdtlU5RcGC3GJ8vjGhQXIWdgaoc1NVH8XF3RIHGuhg)
- [Midjourney Business Breakdown & Founding Story](https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHTXaJ0YazXk6pq45--pfusqT_bogAu98NjBUWEZUYGBNsbtazZWNI35EuBEc2AExLJZpL_KRpp1OYCj4gmdvydB1REvISfz3mgiCWbv-cOxNF6Yk16Q_HIV8bvzhnk7dW9pW3yYA==)