| Media | Tool | Runs on |
|---|---|---|
| Images | ComfyUI + open-weight image models | 8GB VRAM+ |
| Video | ComfyUI-based video pipelines | 24GB+, or 64GB+ unified |
| Music | Local audio models (ACE-Step class) | 12GB+ VRAM |
| Voice | Local TTS/voice-clone (F5-TTS class) | 8GB+ |
| Stems/demix | Demucs-class separation | CPU works, GPU faster |
Our reference machine is a GMKtec EVO-X2 AI (AMD Ryzen AI Max+ 395, 128GB unified memory): image and video generation via a ComfyUI studio, audio via a ROCm-based audio studio, LLMs alongside (GLM 4.7 Flash at 38 tok/s). Unified memory is the cheat code — the model pool isn't capped by a small VRAM card.
| Hardware | Comfortable for |
|---|---|
| 8GB VRAM (any GPU) | SD-class images, voice |
| 12-24GB VRAM | Large image models, short video, music |
| 64GB+ unified (Strix Halo / Apple Max) | Frontier video + audio models, everything at once |
Our pre-configured AI-OS sets up the workstation side (models, runtimes, dashboards), and the Local Media Studio wraps ComfyUI with a character system and MCP server for automation — both in daily studio use.
AI-OS — the workstation stack MUSE Suite (free download) All downloadsIf you generate a handful of images a month, cloud services are simpler and cheaper than a $700+ machine. Local wins on privacy, volume, custom characters, and zero marginal cost — not on convenience.