Threat Advisory

Ollama and Nvidia Triton Inference Server Vulnerabilities Enable Arbitrary Code Execution

Threat: Vulnerability
Targeted Region: Global
Targeted Sector: Technology & IT
Criticality: High
[subscribe_to_unlock_form]


EXECUTIVE SUMMARY:

A set of vulnerabilities affecting popular AI inference platforms, including Nvidia Triton and Ollama, could allow local and remote attackers to disrupt services, bypass authentication, access or copy arbitrary files, and achieve remote code execution against model-serving infrastructure if left unmitigated. Organizations running AI model runners or inference servers should immediately verify and apply vendor updates, restrict network exposure of inference endpoints, enforce strong access controls and authentication, isolate model-serving instances through containerization and network segmentation, harden filesystem and configuration permissions, enable logging and monitoring for anomalous activity, and follow least-privilege principles for any service or user that can modify model configuration or deployment files to reduce the risk of compromise and lateral movement.[/subscribe_to_unlock_form]


EXECUTIVE SUMMARY:

A set of vulnerabilities affecting popular AI inference platforms, including Nvidia Triton and Ollama, could allow local and remote attackers to disrupt services, bypass authentication, access or copy arbitrary files, and achieve remote code execution against model-serving infrastructure if left unmitigated. Organizations running AI model runners or inference servers should immediately verify and apply vendor updates, restrict network exposure of inference endpoints, enforce strong access controls and authentication, isolate model-serving instances through containerization and network segmentation, harden filesystem and configuration permissions, enable logging and monitoring for anomalous activity, and follow least-privilege principles for any service or user that can modify model configuration or deployment files to reduce the risk of compromise and lateral movement.[emaillocker id="1283"]

  • CVE-2024-12886: It is a memory-exhaustion vulnerability in the Ollama model server that can be triggered by a crafted GZIP HTTP response which the server reads without bounds, allowing an attacker to exhaust memory and crash the service. The root cause is the use of io.ReadAll on untrusted response bodies in code paths such as makeRequestWithRetry and getAuthorizationToken. The vulnerability has a CVSS score of 7.5.
  • CVE-2025-51471: It is an authentication-bypass vulnerability in the Ollama model server that allows access to protected functionality without valid credentials. It can be exploited by sending crafted requests that subvert the servers authentication checks due to improper validation in request-handling paths. The vulnerability has a CVSS score of 6.9.
  • CVE‑2025‑48889: It is an arbitrary file‑copy vulnerability in Gradios flagging feature that allows unauthenticated to copy any readable file from the servers filesystem. Although attackers cannot directly read the copied files, they can trigger denial-of-service by copying large files like dev.urandom, filling disk space. The vulnerability has a CVSS score of 7.5.

 

RECOMMENDATION:

  • We strongly recommend you update Ollama to version v0.12.10.

 

REFERENCES:

The following reports contain further technical details:

[/emaillocker]
crossmenu