Skip to main content
July 23, 202618:20

I Sped Up My Local AI 64% with this $100 Gadget!!

By Samuel Gregory

About this video

Large fan (UK): https://amzn.to/4fMFkIR Large fan (US): https://amzn.to/4fej0rB Small fan (UK): https://amzn.to/4wUFW55 Small fan (US): https://amzn.to/4pxCGKf Your high-spec MacBook is lying to you about its speed. In this video, I test how to stop thermal throttling on the M5 Max MacBook Pro using external cooling systems to see if we can truly run massive local AI models at peak performance. Key Takeaways: - Why the 14-inch MacBook Pro throttles during heavy AI workloads. - A comparison between the Base V12 and V10 cooling systems. - How to create a custom 'seal' to increase cooling efficiency. - The data: how I achieved a 64% increase in AI generation speed. - Why local AI is the future for founders who value speed and privacy.

CAD files from this video

Download .zip
  • panel.stepParametric CAD model for editing in Fusion, FreeCAD, etc.
  • panel.stlMesh ready for slicing and 3D printing.
  • panel.pyPython script used to generate the model.

The Era of the Out-of-the-Box Setup is Dead

If you are a founder running a high-spec MacBook Pro, you are likely leaving half of your performance on the table. We buy these machines for the promise of 'unlimited' power, but the laws of physics do not care about your bank balance or your RAM specifications. Thermal throttling is the silent killer of productivity in the age of local AI.

In my latest experiment, I took a top-spec M5 Max MacBook Pro with 128GB of RAM and pushed it to the limit with a 35B local coding model. The results were clear: the 14-inch form factor, while perfect for the nomadic CEO, simply cannot handle the heat of sustained AI workloads.

The Problem: The Portability Tax

We choose the 14-inch model because we value the ability to move between coffee shops, boardrooms, and home offices. However, when you are running local LLMs, that small chassis becomes a furnace. Within minutes, the system throttles the clock speed to prevent damage, effectively turning your £4,000 powerhouse into a mid-range machine.

The Solution: Bespoke Cooling

I tested two cooling systems: the Base V12 and the V10. The goal was to see if we could combine 16-inch performance with 14-inch portability.

The findings were staggering:

  1. The 'Swimming Pool' Effect: High-end coolers create a sealed environment of high-pressure air.
  2. Custom Engineering: I had to fashion a custom 'jig' out of the box itself to ensure the air was directed precisely into the Mac's intakes.
  3. The ROI: A 64% increase in token generation speed.

Why This Matters for Founders

Local AI is about more than just speed: it is about data sovereignty and personal software. When you run your models locally, your IP stays on your hardware. But to make local AI a viable part of your workflow, it has to be fast.

You do not need a Mac Studio to get pro-level performance. You need a strategic approach to your hardware stack. Sometimes, that means spending £98 on a fan and a little bit of time on a DIY seal to unlock the true potential of your silicon.

Keep on vibing.

Transcript

So, I recently bought this M5 Max MacBook Pro with 128 gig of RAM for all of that local AI goodness. It is a 14-inch and I have seen the comments about thermal throttling. I knew that going into it, but I personally just like the 14-inch factor. Thermal throttling is an issue and I know I will not be getting the best out of it when pushing 128 gig on local AI.

I found a cooling system which will hopefully give the 14-inch that extra cooling oomph it needs at my desk. Does it work? Will it make my AI faster? We are putting it to the test so you can combine portability with a cooling powerhouse.

I have not seen anyone do videos on using these for AI specifically. Gamers use them, but no one is doing AI stuff, so we might be a first here. This cooling pad creates a sort of 'swimming pool' effect with the fans. It cost £98, so it is not cheap.

The software only runs on PCs because this is for gamers, and it is not compatible with Mac OS. That is a shame because you lose some advanced functionality. I ended up creating a custom frame out of the box to hold the laptop and create a better seal. This stops the air from escaping the sides and points it directly at the intakes.

I tested a 35B local coding model. Without the fan, it hit 90 degrees almost instantly and then throttled. With the fan and my custom jig, it still hit 90 degrees, but it held there and sustained performance. The results were insane: 2 minutes and 49 seconds compared to over 7 minutes.

I also tested a smaller fan, the V10, but it did not perform as well. After averaging the tests, the bigger fan led to a 64.12% increase in tokens per second. The power gadget graphs show that the clock speed is sustained for much longer with the fan, whereas without it, the GPU throttles significantly.

The verdict? You will definitely get a speed increase. If you want the portability of a 14-inch but the power for heavy AI tasks at home, this is a viable option. It makes you wonder if a Mac Studio is a better choice, but for those of us who travel, this hardware hack is a game changer.