OpenAI has released gpt-oss-120b, a 120-billion parameter language model, but running it on standard consumer hardware reveals significant performance bottlenecks. This release highlights the gap between accessible model weights and the practical requirements for local inference. Readers interested in local AI deployment should note that even quantized versions demand substantial resources.
Experiment shows heterogeneous memory pooling is too slow for practical use
A single user tested the 4-bit quantized version of gpt-oss-120b by combining six disparate devices into a single cluster. The setup included a Samsung Galaxy S24+, an RTX 3060 mini PC with 12GB of VRAM, a Windows notebook with 12GB of RAM, an Intel MacBook Pro, a Mac mini, and an M3 MacBook Pro. This heterogeneous approach attempts to pool memory across iOS, macOS, and Windows environments.
The experiment utilized a technique called pipeline parallelism to distribute the model layers across the available devices. This method reduced the total memory footprint to 47GB, which is below the estimated 60GB to 80GB minimum required for decent performance. The source noted that a 120B LLM typically requires a combination of VRAM and system memory to function effectively.
Despite the successful configuration, the cluster generated output at a speed of just 1.1 tokens per second. This rate is too slow for practical productivity tasks, confirming that the approach is impractical for everyday use. The experiment demonstrates that while memory pooling is technically feasible, it does not solve the latency issues inherent in running large models on low-end hardware.
The test confirms that gpt-oss-120b can run on a mix of consumer devices, but the resulting speed makes it unsuitable for real-time applications. Users seeking to run 120B parameter models locally will likely need dedicated high-end workstations rather than a collection of low-spec machines.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.