Disclaimer: This research was conducted independently. Views expressed are my own and do not represent my employer.
Introduction
I’ve been diving into various protocols here lately during my research and have uncovered another Server-Side Request Forgery vulnerability, this time in the popular local AI model inference engine Ollama. Ollama gives users a quick and free way to run local AI models on their edge devices. A typical use case would be an employee spinning up Ollama on their laptop to run a smaller model, such as Qwen3 or similar, to help automate and augment some of their tasks. What drew me to this target is that it also offers a free way for companies to stand up a chatbot style interface for internal employees, which may leave the APIs exposed, and the fact that the Docker image for Ollama serves on 0.0.0.0 by default with zero authentication on its APIs.
Ollama’s APIs
Ollama exposes various APIs by default
- Pull - Pull a model from an OCI registry
- Push - Push a local model to an OCI registry
- Copy - Copy/rename a model locally
- Delete - Delete a local model
- Tags - List local models
- Chat - Run inference
- Generate - Text completion
Finding the Bug
After I discovered CVE-2026-33990 in Docker’s Model Runner, I knew exactly where I wanted to begin searching in Ollama, so I dove into its OCI registry protocol to look at how it handles both realm headers and redirects. My last finding exploited the fact that there was no validation on the WWW-Authenticate: Realm header, which could force the server to request internal endpoints, and if a regex match for ’token’ matched, would return the retrieved token back to the attacker. Ollama seems to have thought about this exact scenario, as they check if the realm host is the same as the original host before following the URL. The snippet from server/auth.go is below.
if redirectURL.Host != originalHost {
return "", fmt.Errorf("realm host %q does not match original host %q", redirectURL.Host, originalHost)
}
The Redirect
I pivoted to looking at how Ollama handled redirects to see if there was anything we could exploit as a remote attacker. When Ollama pulls a model, it downloads each blob (layer) from the registry. Before actually downloading anything, it resolves a “direct URL”, which is standard OCI pattern where a registry can say “don’t download here, go get this blob from this CDN instead”.
Ollama sends a GET to the registry for the blob. If the registry responds with a 307 redirect to a different hostname, Go’s HTTP client hits the CheckRedirect function, which compares the redirect target hostname with the original hostname. If they don’t match, it returns the http.ErrUseLastResponse, which tells Go’s HTTP client to stop following and return 307. But instead of rejecting the redirect, Ollam areads the Location header and uses that URL as the download target. The responsible code snippet from server/download.go is pictured below.
// before this, code checks if hostnames match, if they don't we hit this code block
// here we check if status code is 200 or 307 , and since our malicious OCI registry returns 307, it extracts location header and uses as download target
if resp.StatusCode != http.StatusTemporaryRedirect && resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("unexpected status code %d", resp.StatusCode)
}
return resp.Location()
So as of now, the flow would be to host a malicious registry that will return a 307 redirect and point at http://some-internal-service, and then hit the /api/pull endpoint and point at our server to trigger SSRF.

Bypassing Hash Verification
But thats blind, and while it is recognized as legitimate vulnerability (could use oracles to map the network without a response), its obviously much more fun if you can actually get the responses back from your forged requests.
When Ollama follows our redirect and fetches from the internal service, it writes the response body to a blob file on disk. Normally, after all blob files are downloaded, Ollama runs verifyBlob() which computes the SHA256 hash of the stored blob and compares it against the digest in the manifest. It would obviously be impossible to compute the SHA256 for the internal service response before we even see it, so the blob fails the check and gets deleted from the host.
But there’s a bug in how Ollama tracks which blobs need verification. The download loop in server/images.go uses a map keyed by digest to track cache hits:
// map to track which blobs to skip
skipVerify := make(map[string]bool)
// iterate layers
for _, layer := range layers {
// downloadBlob checks if the blob already exists on disk, and returns true if so, download otherwise
cacheHit, err := downloadBlob(ctx, downloadOpts{
digest: layer.Digest,
// ...
})
skipVerify[layer.Digest] = cacheHit
}
The vulnerability here is that if we serve a model manifest where config and layer share the same digest, the value at skipVerify[layer.Digest] is marked true after the second loop, since the system thinks it already has this digest (I mean, it does). Meaning we can bypass the hash verification step, meaning our SSRF responses persist on disk.
The Size Problem
So we can redirect Ollama to an internal service, bypass hash verification, and now exfiltrate the response via a call to /api/push. There is just one more blocker - Ollama needs to know how many bytes to read.
Before downloading a blob, Ollama sends a HEAD request to our registry to get the Content-Length needed. This value gets used in the pull functionality to control how many bytes to read in. We can control the HEAD response since its our rogue OCI registry, but we don’t know how large the response from the internal endpoint is going to be.
If Content-Length is set too high, it will read the available response, hit EOF waiting for more bytes, and then fail. These failures take about 10 seconds.
If we set Content-Length less than or equal to the actual response, Ollama will read exactly that number and succeeds. This takes about 1 second.
That timing difference gives us the oracle we need to figure out the necessary Content Length to repsond with in order to caputre the full SSRF response. I was actually excited because this was prime opportunity to use a binary search algo to find the size. The steps are as follows:
- Set Content-Length to 65536. If the pull times out, we know its too big.
- Try 32768. If the pull times out, still too big.
- Try 16384. If it succeeds fast, the size is between 16384 and 32768.
- Keep halving until we find the exact byte length.
Professor Widman would be so proud of me for using my binary search to exfiltrate sensitive credentials :’)
This is actually just the easiest way to get the full SSRF responses written to disk. If the goal was to get data as quickly as possible, we could just start with a smaller value, say 100 bytes and see where that gets us.
Exfiltration
Now that we have full SSRF responses written to disk, there is only two more requests we need to make to exfiltrate the data.
Copy using the /api/copy API with the source set to our rogue registry model name and destination set to a name pointing at our exfiltration server. This just creates a second reference to the same blob locally.
Push using the /api/push with the destination model name. Ollama connects to our exfil server using OCI push protocol.
Proof of Concept
The POC below uses the Ollama Docker image (binds to 0.0.0.0 by default) alongside a Consul instance running in dev mode on the internal Docker network. Consul in dev mode has no ACLs enabled and exposes its full KV store over HTTP, which a realistic setup commonly found in development and staging environments. We use the poc script to enumerate live endpoints based on a target list and then exfiltrate the Consul KV store, which contains database connection strings, Redis auth tokens, JWT signing keys, and OAuth secrets.

SSRF via malicious OCI model pull
Disclosure
Ollama was contacted multiple times with no response.
- CVE: CVE-2026-5530
- PoC: GitHub
Thanks for reading my blog <3