Editor's note: Today we announced that Stealthium has joined the Vultr Cloud Alliance, bringing GPU observability and runtime security to Vultr Cloud GPU. Read the announcement here.
Not long ago, the hardest question in AI infrastructure was whether you could get GPUs at all. Thankfully, there is now a wave of neoclouds committed to providing access to GPUs for their customers. This answers the capacity concern. It does not, however, always answer the confidence concern when teams have limited visibility of how AI moves from pilot to production, or what exactly is happening at runtime. That is why we have joined the Vultr Cloud Alliance.
Security As A Bottleneck
Providers like Vultr have turned the latest AMD and NVIDIA accelerators into something you provision in minutes, across 33 regions, as a virtual machine, bare metal or a self-service cluster.
Capacity is no longer the bottleneck. Security still is.
The organisations now arriving at neoclouds are not only start-ups racing to train. Increasingly they are enterprises and regulated operators moving production workloads, and they bring a decade of hyperscale cloud governance expectations with them. Their questions are can we prove what is running on purchased capacity, how can we ensure our most sensitive workloads are secure, and what is our burden of a shared responsibility model?
Sign-Offs Between Pilot And Production
We have found between a successful pilot and a production AI service, there can be three different set of sign-offs.
Infrastructure and finance: Are we getting what we pay for? Where are we wasting capacity, are any of our cards degrading, and why did that training run fail at hour forty?
Customers and regulators: Can you demonstrate, rather than assert, how our data and models were handled?
Security and risk: Is this workload isolated from whoever else shares the hardware? Would we know if something was running on this accelerator that should not be? Can we show an auditor evidence rather than a tenancy diagram?
These approvals are where AI adoption stalls.
A recent survey (EMA with Protegrity) found that 82.9% of organisations had seen AI projects delayed before production by security or compliance reviews, and 81.6% had deployed AI in a reduced form because of security concerns.
Observing GPUs In Real-Time
The telemetry that tells a security team a GPU is running something it should not, is the same telemetry that tells an infrastructure team a GPU is wasted, throttling or failing.
Cryptojacking and a forgotten job both burn capacity you are paying for. Similarly, an unauthorised inference workload and a runaway training loop look like load nobody planned. A card drifting towards failure and a card under attack both show up first as behaviour that departs from the workload's normal pattern.
Without runtime visibility, you cannot tell these apart. With it, the same signal serves both infrastructure and security teams.
Operational waste is not negligible. In just one view, Cast AI's analysis of tens of thousands of Kubernetes clusters put average GPU utilisation at 5% before optimisation.
In this environment, runtime security and observability, rather than being an added cost, can be the operating layer for AI adoption that offsets its own cost. CFOs wanting utilisation evidence and the CISOs wanting isolation evidence can get both, with the same signal. Stealthium provides that.
AI Evidence Needs To Exist
Every other layer of the stack can already answer these questions. Networks, hosts, containers and applications all have mature telemetry, detections and audit trails. GPUs, aka AI accelerated compute, the most expensive and most privileged component in the environment, is the exception.
Traditional monitoring can tell you a GPU is busy. It cannot tell you what it is busy doing. Security tooling built for the CPU stops at the driver. Below that line is where the workload executes, where memory is shared and where isolation is tested, and none of it produces a signal a security team can act on. If you follow our recent research, whether it's Behind Bars or GPUBreach or Januscape, you will know that the isolation boundaries multi-tenant deployments depend on can be compromised - a class of attack that executes on the GPU and never reaches CPU-focused tooling. If you have read our analysis of ShadowRay, you will know these environments are being attacked in the wild.
Composable Compute Needs Composable Confidence
Vultr delivers a range of choice for their customers. Customers can pick their cloud and silicon for cost, availability and fit, and change their minds as workloads change. Vultr offers their customers composable compute, and we are thrilled to add to the partnership with composable confidence.
Stealthium deploys as a single agent per host on Vultr Cloud GPU instances, with no changes to the Vultr platform. Teams add GPU-level telemetry, workload context, health monitoring and runtime threat detection to the environments that need it, when they need it, whether it is their first deployment or their fiftieth.
It also makes shared responsibility workable at the GPU layer. The provider secures and operates the platform. The customer owns what runs on it.
We announced this partnership, early this morning at the Cloud & AI Infrastructure Singapore. It was a fitting venue, as across Asia Pacific, regional and sovereign AI capacity is being built at pace, and the regulated operators it serves need to see and prove what runs on it.
From Access To Confidence
The first phase of the neocloud era was about access. The next is about confidence: proving to a board, an auditor or a customer what your AI compute is doing, and knowing you are getting what you pay for. Organisations that can produce that evidence will move workloads into production. Those that cannot will keep them in pilot, however much capacity they have reserved.
Capacity gets you started. Confidence gets you to production.
