Skip to content
GMKtec|Enthusiast tier

GMKtec EVO-X2 (128GB)

Specs not independently verified
Starting at$3,500

Design score

8.0/10

How we score →
Design & Build8.0
GMKtec EVO-X2 (128GB)
ChipAMD Ryzen AI Max+ 395
RunsLlama 3.3 70B, Mistral Small 3.1
Throughput15-25 tok/s
Unified memory128 GB

Verdict

Enthusiast-tier mini workstation: AMD Ryzen AI Max+ 395 with 128 GB of unified memory. Runs Llama 3.3 70B and Mistral Small 3.1 at roughly 15-25 tokens/s — 70B-class models without a discrete GPU, in a small-form-factor chassis.

Where it wins, where it doesn't

Pros

  • 128GB massive memory pool
  • Ultra-compact form factor
  • Dual NVMe storage slots

Cons

  • APU graphics bottleneck token generation speed
  • Audible fan noise under sustained load
Ideal forHigh-capacity context retrievalCPU-based LLM inference

Specifications

ChipAMD Ryzen AI Max+ 395
RunsLlama 3.3 70B, Mistral Small 3.1
Throughput15-25 tok/s
Unified memory128 GB

Editorial note

Packing 128GB of RAM into a mini-PC chassis enables running massive models that simply will not fit in standard consumer VRAM. While inference speed via CPU/APU is slower than discrete GPUs, the sheer memory capacity unlocks use cases previously restricted to cloud servers.

In-Depth Review

The EVO-X2 packs an AMD Ryzen AI Max+ 395 with 128GB of unified memory into a mini-PC chassis — enough to run Llama 3.3 70B and Mistral Small 3.1 at roughly 15–25 tokens/second without a discrete GPU.

What the memory unlocks

Models in the 70B class simply do not fit in standard consumer VRAM. 128GB addressable by the APU makes them loadable on a machine you can hold in one hand, drawing a fraction of the power a multi-GPU rig would. Dual NVMe slots handle the model library.

The trade-off is speed, as always with this architecture

The APU graphics bottleneck token generation — 15–25 tok/s is usable for interactive chat and retrieval, not for high-throughput serving or batch work. And the fan is audible under sustained load.

Who should buy

High-capacity context-retrieval workloads, CPU/APU-based LLM inference, and anyone who needs 70B-class capability in a small, quiet-ish box. If your target models fit in 16–24GB of VRAM, a discrete GPU runs them several times faster for less money.

Frequently Asked Questions

Who is GMKtec EVO-X2 (128GB) for?
GMKtec EVO-X2 (128GB) is a fit for high-capacity context retrieval and CPU-based LLM inference.
What are the drawbacks of GMKtec EVO-X2 (128GB)?
The trade-offs we record are: APU graphics bottleneck token generation speed and Audible fan noise under sustained load.
What does GMKtec EVO-X2 (128GB) do well?
128GB massive memory pool, Ultra-compact form factor and Dual NVMe storage slots.
How much does GMKtec EVO-X2 (128GB) cost?
$3,500, as recorded in this index. Prices move; check the retailer for the current figure.
What is the Fathom Layer design score for GMKtec EVO-X2 (128GB)?
GMKtec EVO-X2 (128GB) scores 8.0 out of 10. The score is assigned by a human against published criteria and placement is never sold — see the methodology page for how it is built.

Further reading

Featured badge

Building this product? Add the badge to your site to show it’s in the index.

<a href="https://fathomlayer.com/compute/local-ai-workstations/gmktec-evo-x2-128gb" target="_blank" rel="noopener noreferrer"><img src="https://fathomlayer.com/fathom-badge.svg" alt="Featured on Fathom Layer" width="250" height="54" /></a>