Every model below is installed and in daily studio use — this is not a listicle. Sizes are real disk footprints from our machines; speed notes are measured, not vendor claims. Reference hardware: GMKtec EVO-X2 AI (AMD Ryzen AI Max+ 395, 128GB unified memory) plus an 8060S-class iGPU workstation. The whole stack runs through AI-OS + ComfyUI + Ollama, fully offline.
| Model | Size | Use it for | Measured notes |
|---|---|---|---|
| GLM 4.7 Flash (30B-A3B MoE) | 19GB | Daily driver: chat, analysis, agents | 38 tok/s on 128GB unified; :32k variant for long context |
| Qwen3.8 (27B) | 17GB | General + coding | ~40 tok/s class on Strix Halo |
| Ornith 1.5 (35B-A3B MoE / 9B) | 22GB / 6.7GB | Strong MoE alternative; 9B fits smaller machines | New addition, 32k context variants |
| Qwen2.5-Coder (32B/14B/7B) | 19/9/4.7GB | Specialized code generation | The 7B is the best coding-per-GB in our set |
| DeepSeek-R1 7B | 4.7GB | Reasoning chains | Slow but deliberate |
| Llama 3.2 3B / Gemma2 2B | 2GB/1.6GB | Edge tasks, classification | Instant load on anything |
| nomic-embed-text | 274MB | Embeddings / RAG | Powers local search |
| Model | Size | Use it for | Measured notes |
|---|---|---|---|
| z-Image Turbo | 12GB | Fast image generation — our house generator | ~26s per 1216x832 on the mini PC; generated every thumbnail on qalarc.com |
| FLUX.1 Krea-dev (fp8) | 12GB | Highest-quality creative images | Slower than turbo, noticeably better composition |
| Wan 2.2 ti2v 5B | 9.4GB | Fast text-to-video | The practical entry point for local video |
| Wan 2.2 Animate 14B | 17GB | Character animation from reference | 64GB+ recommended |
| Wan 2.2 S2V 14B | 16GB | Speech-driven video | Audio-to-talking-head pipeline |
| Wan 2.1 flf2v 720p 14B | 16GB | First/last-frame video interpolation | Keyframe-driven clips |
| LTX 2.3 22B (+distilled LoRA) | VAEs 2GB + LoRA 7GB | Video with audio track | One-pass video+audio generation |
| SAM 2.1 Hiera | ~400MB | Segmentation | Video object masking |
| RIFE v4.26 | small | Frame interpolation | Smooths generated video |
LoRA stack worth knowing: cyberrealistic-z-Image (photoreal), unstableRevolution V3 (SD-class creative), famegrid (house style). Text encoders that matter: Qwen3-4B and Qwen2.5-VL-7B (multimodal understanding), T5-XXL / UMT5-XXL (classic conditioning).
| Model | Use it for | Notes |
|---|---|---|
| ACE-Step | Full music tracks from text prompts | Genre-accurate, runs locally on GPU |
| Stable Audio 3 | General audio synthesis, SFX | Fast iteration |
| F5-TTS | Voice cloning from short samples | Best quality-per-ease in our set |
| GPT-SoVITS | Multilingual voice | Cross-lingual cloning |
| MelBand Roformer | Stem separation | State-of-the-art demixing |
| Vocos | Neural vocoder (mel to audio) | The speed layer under TTS pipelines |
| Wav2Vec2 (large-xlsr) | Speech recognition features | Feeds audio-driven pipelines |
Unified memory changed the math: a 128GB mini PC runs 30B MoE language models and 14B video models that would need an expensive discrete-GPU rig otherwise. The pattern we would recommend to anyone: one Strix-Halo-class machine, Ollama for text, ComfyUI for vision, a dedicated audio stack — everything offline, zero marginal cost per generation.
AI-OS — the workstation stack Local media guide The mini PC writeupOccasional use (a few images a month) is cheaper on cloud services. Frontier-class quality still lives behind APIs. Local wins on volume, privacy, custom characters, and cost-per-generation at scale.