# LLM Inference Machine for $300

DevFeed: [LLM Inference Machine for $300](<https://devfeed.tech/articles/llm-inference-machine-for-300-27454.md>)

Original publisher: [Read original article](<https://ariya.io/2024/12/llm-inference-machine-for-300/>)

Published: 2024-12-28T04:17:14Z

Content type: article

Language: en

Sources: [Ariya Hidayat](<https://devfeed.tech/sources/ariya-hidayat.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [3d-printed](<https://devfeed.tech/tags/3d-printed.md>), [amd](<https://devfeed.tech/tags/amd.md>), [cost](<https://devfeed.tech/tags/cost.md>), [ddr4](<https://devfeed.tech/tags/ddr4.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [memory](<https://devfeed.tech/tags/memory.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [ssd](<https://devfeed.tech/tags/ssd.md>), [threads](<https://devfeed.tech/tags/threads.md>), [vision](<https://devfeed.tech/tags/vision.md>)

## AI overview

A $300 used-hardware build runs quantized LLMs including Qwen-2.5 32B, Llama-3.1 8B, and Llama-3.2 Vision 11B. It uses an NVIDIA Tesla M40 with 24GB of VRAM and achieves model-dependent speeds from 7 to 47 tokens per second. The article compares its performance, cost, and upgrade options with newer GPUs and Apple Silicon.

## Source excerpt

You can absolutely run Qwen-2.5 32B. And of course, Llama-3.1 8B and Llama-3.2 Vision 11B are no problem at all.