Skip to content

ToshLLM runs open AI models locally on Intel Macs with AMD graphics, including old Mac Pros

ToshLLM is a free, open-source app that runs open language models on Intel Macs with AMD graphics, a class of Mac most local AI tools leave behind.

Pinkesh Gajera4 min read

ToshLLM is a free, open-source Mac app that runs open-weight AI language models entirely on Intel Macs with AMD graphics cards.

AppleInsider pointed Intel Mac Pro owners to it on 9 October. The project is published on GitHub under the GPL-3.0 licence, by a developer who goes by engeldlgado; the repository was created in June 2026, is labelled beta, and released version 0.87.19 on 9 October. Its pitch is the one AppleInsider quotes: local models with no cloud and no bill per token.

Which Macs can run ToshLLM?

The project's README lists three requirements: macOS 14 or later, an Intel Mac with an AMD graphics card that supports Metal, and at least 16GB of memory, with 32GB recommended for the largest models it suggests. External AMD cards in an eGPU enclosure are supported as well as internal ones.

Older machines get a separate download. The normal build needs the AVX2 instructions, and the README says it crashes on launch on the 2010 to 2012 Mac Pro and other pre-2013 Xeon Macs, which should use the build whose name ends in noavx2. From version 0.86.5 the app is signed with an Apple Developer ID and notarised by Apple, so it opens without a security warning; earlier versions needed Open Anyway in Privacy and Security. Everything it needs ships inside the app, with no Homebrew or Python to install.

What the app does

According to the README, the standard local AI engines produce corrupted output on AMD graphics cards in Macs, and read model data over PCIe far more slowly than they could. ToshLLM bundles llama.cpp, the widely used open-source engine, with AMD-specific changes, including a Flash Attention kernel so that attention runs on the graphics card rather than falling back to the processor. Around it sits a SwiftUI app with chat, a model manager that estimates memory use for the Mac it is running on, Hugging Face browsing, image input, a beta image generator and an OpenAI-compatible local server at port 8080 for other apps to use.

How fast is it?

The developer's own figures, measured on a Radeon RX 6700 XT with 12GB, are below. Prompt speed is how fast the model reads what you give it; generation is how fast it writes. Nobody independent has published tests yet.

Model

Prompt, tokens a second

Generation, tokens a second

Llama 3.2 1B

5770

254

Qwen3 8B

851

61

gpt-oss 20B

1305

94

Qwen3.6 35B-A3B

475

29

For Qwen3 8B, the README puts unmodified llama.cpp on the same card at 0.6 to 2.6 tokens a second. It also sets its gpt-oss 20B result against figures posted by Apple silicon owners in a llama.cpp discussion, and finds generation about level with an M4 Max at roughly 95 tokens a second. The developer adds the caveats: the model files differ slightly, and one of the Apple silicon results was throttled by heat.

Our take

Intel Macs have been left out of Apple's AI story from the start. Apple Intelligence has only ever run on Apple silicon, and Apple has stopped making macOS for Intel altogether, which is why owners of older machines have turned to tools like OpenCore Legacy Patcher, the route to Tahoe for Intel Macs Apple no longer supports. That patcher matters here too: the 2010 to 2012 Mac Pro the noavx2 build is aimed at cannot run macOS 14 without it.

The headline numbers need reading with care. The RX 6700 XT is not a card Apple ever put in a Mac, and the README's notes on getting RDNA 2 cards like it working are written for Hackintosh builds. A 2019 Mac Pro with Apple's own MPX graphics modules, or a 16-inch MacBook Pro with 4GB or 8GB of graphics memory, is different hardware, and a laptop will fit far smaller models. The app's memory estimates for your own machine are a better guide than the chart.

It belongs to a pattern we keep seeing, of individual developers doing for Intel Macs what Apple will not, from NullMoth's Nvidia driver for macOS Sequoia to eGPU projects that only work because Apple still supports external graphics on Intel Macs alone. APPDOOK's view is that ToshLLM is worth a try for anyone with a Mac Pro or an AMD eGPU who wants a private local model, on the understanding that it is one developer's beta. The upside of Apple freezing Intel Macs at Tahoe is that the target has stopped moving, which makes a project like this easier to keep working than it would have been a year ago.

Sources

  1. ToshLLM: Local AI for Intel Macs with AMD GPUsToshLLM, 2026-10-09
  2. toshllm: Run large language models locally on Intel Macs with AMD GPUsGitHub, 2026-10-09
  3. ToshLLM v0.87.19 release notesGitHub, 2026-10-09
  4. You can run AI models on that Intel Mac ProAppleInsider, Mike Wuerthele, 2026-10-09

Reporting and images linked above belong to their respective publishers and are shown from their own servers. The analysis here is our own.

MacIntel MacMac ProAIOpen source

Share this article

Keep reading