The upstream Linux amdgpu driver, unmodified, running on macOS.
AMD Radeon GPUs over Thunderbolt on Apple Silicon Macs, driven by the same
amdgpu + amdkfd code Linux uses, inside a DriverKit system extension.
mac_linuxgpu · amdgpu_mtopg · LemonSeed Engine
Apple Silicon Macs have Thunderbolt 5 but no AMD GPU driver. mac_linuxgpu
runs the real Linux kernel driver instead of reimplementing one. It compiles
the upstream amdgpu, amdkfd, TTM, DRM core and GPU scheduler sources
(468 files from a pinned Linux commit, unmodified) against linuxu, a
userspace implementation of the Linux kernel APIs those sources need, and
hosts the result in a PCIDriverKit extension.
The GPU initializes the way it does on Linux: IP discovery, PSP, SMU, GMC, GFX12, SDMA and MES. Compute works the way it does on Linux too. Every client is a KFD process with its own GPU address space, and its HSA queues are MES-scheduled user queues. The goal is that GPU software written for Linux/ROCm runs on a Mac with nothing Mac-specific.
LemonSeed Engine, Qwen3.8-27B Q4 with DFlash2 speculative decoding, on an AMD Radeon AI PRO R9700 over Thunderbolt 5 (MacBook Pro, Apple M5 Max), running on this driver:
| Measured | |
|---|---|
| Decode, warm (DFlash2, 74% draft acceptance) | 43.3 tok/s |
| Prefill, 73-token prompt, warm | 60.0 tok/s |
| Operations that fell back to the CPU | 0 |
Speculative decoding speed scales with how many drafted tokens are accepted, which depends on the text being generated. With the measured 80 ms per speculative step, the same setup gives:
| Draft acceptance | 74% (measured) | 80% | 90% | 96% |
|---|---|---|---|---|
| Decode tok/s | 43 | ~52 | ~71 | ~87 |
These are first-run numbers. Nothing in the driver has been tuned for throughput yet: each doorbell write and each completion poll still crosses into the driver. Mapping doorbells into the client and using KFD events are the next steps.
- Upstream driver bring-up. The unmodified
amdgpu_pci_probecompletes on the R9700, and every firmware image is loaded on demand by name. - KFD compute sessions. There is one KFD process per client, each with its
own GPUVM. Buffers go through
ALLOC_MEMORY_OF_GPU/MAP_MEMORY_TO_GPU, and queues go throughCREATE_QUEUE→ MESADD_QUEUE, exactly as on Linux. - HSA runtime.
libhsa-runtime64.dylibis installed with the driver, in/usr/local/liband/Library/MacAMDGPU/runtime. HRX/Loom-based software such as LemonSeed Engine finds it without configuration. - Clean lifecycle. Clients can come and go. A session closes through upstream removal, interrupt drain and endpoint isolation, and the driver stays reusable.
- Observability. A read-only observer reads the driver's retained log and
session state at any time (
scripts/read-driver-log.py). While compute runs, it also reads upstream's own telemetry: sysfs attributes such asgpu_busy_percent,mem_busy_percent,gpu_metricsand hwmon, andAMDGPU_INFO. That's the same dataamdgpu_topreads on Linux, and what amdgpu_mtopg displays. Events that matter after a hang or a reboot (driver start, probe result, GPU ring timeouts and device dumps, errors from upstream, removal, quarantine, session close) also go to the macOS unified log, prefixedmac.linuxgpu: EVENT; routine lines stay in the retained log only. Read them, from any boot the log still holds, withlog show --last 1d --predicate 'eventMessage BEGINSWITH "mac.linuxgpu: EVENT"'(DriverKit logs through the kernel, so they show askernel). - Vulkan (in progress). Mesa's RADV builds for macOS on the same upstream amdgpu DRM interface, through a libdrm over the driver's Linux-file transport, and runs offline on the test suite's software GPU. llama.cpp's Vulkan backend builds against it. See vulkan/README.md.
- Nothing hard-coded to one GPU. The driver matches by AMD vendor ID and display/compute PCI class, and upstream's PCI ID table and IP discovery decide what is supported. Queue layout, MQDs, firmware names, ISA and limits all come from the device.
amdgpu_mtopg monitoring the
R9700 through this driver while LemonSeed Engine generates text: GPU load from
GRBM_STATUS samples, SMU clocks and power, hwmon temperatures and fan, and
VRAM in use by the model.
your app (LemonSeed, HRX/Loom, other HSA clients)
│ HSA API
libhsa-runtime64.dylib installed with the driver
│ IOKit user client
┌──────┴──────────────────── DriverKit extension ────────────────────┐
│ per-client KFD process: /dev/kfd + DRM render node, in-process │
│ upstream amdkfd + amdgpu + TTM + DRM + drm_sched (unmodified) │
│ linuxu: the Linux kernel API in userspace │
│ workqueues · timers · RCU · fences · mm/VMA · mmu notifiers · │
│ sysfs · firmware loader · DMA through DART · MMIO · MSI-X │
└──────┬─────────────────────────────────────────────────────────────┘
│ PCIDriverKit: BARs, DMA, interrupts
AMD GPU over Thunderbolt
- Upstream stays upstream. The Linux sources come from a git submodule pinned to one Linux commit. The build carries one declared patch (a 3-line buddy-allocator rollback fix) and two declared compile-time interventions. A verifier checks that nothing else differs from upstream.
- linuxu is where the porting happens. It provides Linux semantics for locking, workqueues, timers, RCU, DMA fences, memory management, notifiers, sysfs and the firmware loader, and the PCI, DMA, MMIO and interrupt seams over DriverKit.
- Firmware is fetched from linux-firmware at build time, verified against a lock file, and served to the driver on demand by name.
- An Apple Silicon Mac and an AMD GPU in a Thunderbolt enclosure
- Xcode with the DriverKit SDK (
xcode-selectpointing at Xcode) python3,gitandcurl(from the Xcode command line tools)cmakefor the HSA runtime (brew install cmake)ripgrepfor some test scripts (brew install ripgrep)- To sign and install the driver: an Apple developer team with the DriverKit PCI and system-extension entitlements, and matching provisioning profiles
There are two ways. Both end in the same tree.
Option 1: the setup script (recommended). It downloads only what the build needs.
git clone https://github.com/lemonade-sdk/mac_linuxgpu.git
cd mac_linuxgpu
scripts/bootstrap.shscripts/bootstrap.sh does all one-time setup and is safe to re-run:
- checks the tools listed above;
- fetches the pinned Linux kernel into
third_party/linuxas a shallow, partial, sparse checkout of only the paths the build reads (about 70 MB downloaded, 600 MB on disk); - applies the patches in
patches/linux/; - downloads the locked linux-firmware files into
build/firmware/and verifies their hashes; - verifies the result and writes a stamp under
build/setup/.
Option 2: plain git.
git clone --recurse-submodules https://github.com/lemonade-sdk/mac_linuxgpu.git
# or, in an existing clone:
git submodule update --initThis works, but git checks out the entire kernel. Measured: about 3.3 GB
downloaded and 1.7 GB on disk, against about 70 MB with the setup script. The
first make (or scripts/bootstrap.sh) then converts the checkout to the
sparse set the build uses, applies the patches and fetches the firmware.
Mesa (third_party/mesa) and llama.cpp (third_party/llama.cpp) are
optional submodules for the Vulkan build only. They are marked
update = none, so neither option fetches them; make radv and
make llama-vulkan do (scripts/bootstrap.sh --mesa-only, --llama-only).
To keep the kernel at the right commit when you pull updates, use
git pull --recurse-submodules, or set it once with:
git config submodule.recurse trueEither way, you don't have to run setup by hand. Every make target
(except clean, distclean, hsa and hsa-test) checks the setup stamp. If
the kernel checkout is missing, at the wrong commit, or unpatched, or the
firmware is missing, make runs scripts/bootstrap.sh first.
scripts/activate.sh and the Xcode project's "Build Linux KMD" phase both go
through make, so a fresh clone builds from any entry point.
make lib-dext # the DriverKit build of the driver library
make test # the offline test suite (no GPU needed)
scripts/activate.sh # build, sign, install and activate the driver| Command | Result |
|---|---|
make / make lib |
build/libmacamgdu.a, host platform, used by the tests |
make lib-dext |
build-dk/libmacamgdu-dk.a, DriverKit platform |
make dext |
links the dext with make (unsigned unless DIDENTITY is set) |
make test |
the offline test suite |
make hsa, make hsa-test |
the HSA runtime (build/hsa) and its unit tests |
make libdrm-mlg |
libdrm and libdrm_amdgpu over the Linux-file transport (build/libdrm-mlg) |
make radv |
Mesa's RADV Vulkan driver and its ICD (build/radv/install) |
make llama-vulkan |
llama.cpp with its Vulkan backend (build/llama.cpp/bin) |
make verify-source |
checks the upstream tree against the pin and patch set |
make clean |
removes build outputs, keeps setup and firmware |
make distclean |
removes build/ and build-dk/ |
The Xcode project mac_linuxgpu.xcodeproj builds the signed dext
(MacLinuxGPU) and the host app (MacLinuxGPUHost). Its first build phase
runs make lib-dext, and the app bundles build/firmware/amdgpu.
scripts/build-installer-dmg.sh produces a signed disk image whose app
installs the driver, the HSA runtime and the firmware.
scripts/activate.sh builds everything, signs with Xcode automatic signing or
locally approved provisioning profiles, installs the app to /Applications,
installs the firmware and HSA runtime, and waits for the driver to attach. Set
XCODE_TEAM_ID to sign with your own team. See scripts/activate.sh --help.
third_party/linuxis a git submodule of torvalds/linux pinned to1f63dd8ca0dc05a8272bb8155f643c691d29bb11. Its sparse checkout covers the amdgpu tree, the DRM core, TTM, the scheduler, a fewlib/helpers andinclude/. Paths that differ only in letter case are excluded, so the checkout stays clean on case-insensitive APFS.mk/upstream_sources.mklists every upstream.cfile the build compiles. Nothing is globbed. The include paths inmk/host_clang.mkpoint straight into the submodule.patches/linux/*.patchare applied to the submodule working tree by the setup script, never committed inside it. Re-applying is detected and skipped. To change a patch, edit it, update itssha256inpatches/manifest.json, and runscripts/bootstrap.sh --reset-linux.scripts/verify-upstream.pyruns asmake verify-sourceand before everylib-dext. It checks that:- the submodule is at the pin;
- the sparse set matches and has no case collisions;
- exactly the declared patches are applied and nothing else in the submodule is modified;
- the few verbatim upstream headers kept in
linuxu/headersstill match; - the compile-time interventions match the declared list.
- To move to a newer kernel: update the submodule commit,
linux.pininpatches/manifest.jsonandmk/upstream_sources.mk, then runscripts/bootstrap.sh.
GPU firmware is not stored in this repository. firmware/firmware.lock pins
a linux-firmware release and lists each needed file with its SHA-256 and
size. scripts/fetch-firmware.sh downloads them from
gitlab.com/kernel-firmware/linux-firmware, falling back to the
git.kernel.org mirror, and verifies every hash.
scripts/fetch-firmware.sh --add amdgpu/<file>.binfetches more files at the pinned release and adds them to the lock.scripts/fetch-firmware.sh --allfetches the whole amdgpu directory, for GPUs beyond the ones in the lock. Bundle it withLINUX_FIRMWARE_DIR=build/firmware/full scripts/activate.sh.
At run time the driver asks for firmware by name. The installed host
component serves it from
/Library/Application Support/MacLinuxGPU/firmware/amdgpu.
dext/ DriverKit extension sources, Info.plist, entitlements
host/ host app: installer, activation, firmware servicer, CLI
hsa/ userspace HSA runtime (CMake) and its tests
libmlg_drm/ the Linux-file client library, and libdrm over it (libdrm/)
vulkan/ the Vulkan path: RADV notes, offline test
linuxu/headers/ Linux kernel API headers for userspace
linuxu/src/ their implementation: memory, locking, PCI, DMA, ...
linuxu/tests/ unit and integration tests
mk/ build configuration: CONFIG table, flags, source list
patches/ patches/linux/*.patch and the upstream manifest
firmware/ firmware.lock (the files themselves are fetched)
scripts/ setup, firmware fetch, verifier, installers, tests
tools/ CONFIG table consistency checks
third_party/linux pinned upstream Linux (submodule)
third_party/mesa pinned Mesa for RADV (optional submodule)
third_party/llama.cpp pinned llama.cpp (optional submodule)
- Alpha. Tested on one GPU, the Radeon AI PRO R9700 (
gfx1201), over Thunderbolt 5. Other AMD GPUs supported by upstreamamdgpushould match and probe, but are untested. - Don't kill the driver process. When a DriverKit driver dies while it
owns a PCI device, macOS runs PCI crash recovery, and that can panic the
machine (
IOPCIFamily). Ifscripts/read-driver-log.pyreports a quarantined session, restart the Mac instead. - Upgrades hand the GPU over explicitly. macOS does not stop a running
driver extension when a new version is activated: the old one stays
attached (
systemextensionsctl listshows itterminating for upgrade via delegate) and the new one attaches only once every old instance is gone. The installer therefore asks the running driver to close its session the normal way before it requests the replacement, and to terminate itself once macOS has accepted it (the Retire selector,dext/sources/session_state.h). It reports apps that still hold a session and a quarantine that needs a restart. To finish a stuck upgrade, quit the apps it lists and runMacLinuxGPUHost retire-previous(--forcecloses their sessions);MacLinuxGPUHost instancesshows what is attached. Drivers up to 0.1.129 (build 233) have no Retire: from 0.1.128 (232) they leave cleanly when the GPU enclosure is switched off; older ones need a restart. - Compute only. The display stack is not built. Vulkan has no presentation yet (headless WSI only).
- No GPU reset recovery.
- Sleep loses device memory. When the host sleeps, the Thunderbolt link
goes down and the GPU is reset, so the driver closes the compute session
through its normal close path before the sleep, and clients reload after
wake (the HSA runtime reports
MAC_HSA_STATUS_DEVICE_LOST). Upstream's only suspend path for a discrete GPU evicts all of VRAM to system memory, which doesn't fit a large model on the iPad. A low-power period without host sleep keeps VRAM: compute is quiesced through upstream KFD suspend and resumes where it left off (mac_hsa_agent_prepare_low_power/mac_hsa_agent_resume, orMacLinuxGPUHostApp power-watchon a Mac). Seedext/sources/power_state.h. - One GPU per driver instance.
- No PCIe atomics over Thunderbolt. Linux has the same limit with this card over Thunderbolt, and upstream's non-atomic firmware path is used.
- amdgpu_mtopg: a live GPU monitor for macOS that reads this driver's telemetry. A signed and notarized download is on its releases page.
- LemonSeed Engine: LLM inference on AMD GPUs through HRX/Loom. On macOS it runs on this driver and uses the HSA runtime the driver installs.
Original code in this repository is available under either the
MIT license or the GNU GPL v2, at your
option (SPDX-License-Identifier: MIT OR GPL-2.0-only); see
LICENSE. Upstream Linux code in third_party/linux, and files in
linuxu/ that state a Linux-derived license, keep their own licenses.
Firmware fetched at build time is distributed under its own terms (see its
WHENCE file) and is not part of this repository.
This driver runs the stack behind these results. They are LemonSeed Engine's HumanEval+ runs of Qwen3.8-27B Q4 on the AMD R9700 with HRX/Loom: 1,368 completed generations at standard, 16K and 32K context. They were measured through the earlier mac_amdgpu driver path, and they are the target for this driver.

