PARTNERSHIPSSECURITYAI

Across The AI Accelerator-verse: A Need for Security and Observability

JUL 2026By Ahmed Shosha, CEO & Founder
Across The AI Accelerator-verse: A Need for Security and Observability

As the demand and potential of AI scales exponentially, the infrastructure landscape is both fragmenting and flourishing. New neo-cloud providers and silicon providers are reestablishing a multi-cloud, multi-accelerator wave. This evolution of compute has not had a matching evolution of security. For many, their AI accelerated compute is without runtime security and observability. That gap is why the multi-accelerator era needs a control plane, and why we are building one with Tenstorrent.

Editor's note: Today we announced a partnership with Tenstorrent to make AI accelerated compute observable and secure. Read the announcement here.

The breadth and access to AI accelerated compute is radically changing. It's truly exciting, and an opportunity for all of us building on and with AI. It's a new multi-cloud, multi-accelerator era, good for everyone, except, the people responsible for securing, verifying and controlling this new infrastructure.

Respectfully, it's taken both enterprises and the security industry years to navigate the original multi-cloud movement and establish mature controls that enable a shared model of responsibility for security. This new wave is larger and moving faster. Neoclouds are standing up dedicated AI capacity, and sovereign programmes are constructing national AI factories on their own terms, on their own soil. Underneath and alongside, silicon is evolving, with custom training chips, dedicated inference accelerators, RISC-V designs, and whole new architectures.

Put shortly, enterprises now have access to a portfolio of accelerators chosen for cost, availability, sovereignty, and fit. Security is not currently part of that portfolio.

Every layer of modern infrastructure is observable and has security controls, except for AI accelerated compute. The most valuable and most privileged layer is the least instrumented, with the tooling meant to watch it stopping at the driver. Below that line the runtime goes dark. The runtime where workloads execute, where isolation between tenants is tested, and where any corruption does its work. Security teams have no alerts, no detections, and no signals tied to what is actually executing on the accelerator. Infrastructure teams are paying for compute they cannot monitor. Leadership cannot validate risk in the most critical compute layer they own. This is blindness across the multi-cloud, multi-accelerator-verse.

More environments simply multiply that risk. Every new architecture is an adopted blind spot, with its own runtime, its own failure modes, and its own attack surface that a CPU-centric stack cannot see. For security tools without instrumentation below the CPU, a legitimate training job and an attacker draining your accelerator memory look identical. Add a second accelerator, or a second neocloud, and you now have two places where that's true.

For those thinking this is a thought experiment, the unfortunate reality is that the security bill always comes due. We are closely monitoring and researching this threat landscape. Our work on GPUBreach showed that rowhammer techniques can be turned against GPU memory through NVIDIA's unified virtual memory, corrupting data across the isolation boundary that every multi-tenant GPU deployment depends on. GPUBreach is just one attack against just one vulnerability class. The telemetry we collect to catch it is the same telemetry that surfaces cryptojacking, unauthorised inference, model weight exfiltration, and cross-tenant memory side channels. These aren't theoretical, they're happening, and they're invisible to traditional security tooling. It isn't a gap in signatures or tuning, rather an architectural failure. Your existing stack has no instrumentation below the CPU.

When I talk with teams deploying on this infrastructure, they're faced with uncomfortable questions. Can you prove that this training run was isolated from the tenant sharing that machine? Can you distinguish genuine training and inference load from waste? Can you detect tampering and unhealthy accelerators before failures cascade? Their answers are so often policy documents, tenancy diagrams and unverified assurances. No real evidence, and nothing for a board, an auditor, or a customer who has just asked you to demonstrate rather than assert.

The answer is observability and security that live at the runtime layer, the layer closest to the silicon, treating the accelerator as something to be seen, proven, and controlled regardless of who manufactured it. Trustworthy AI requires security, control and observability that extend below the CPU fold and beyond the driver. Sovereign AI needs sovereign trust, and you cannot claim control over infrastructure you cannot see at runtime.

Thankfully, Stealthium are not alone in this mission, and our announcement today of our partnership with Tenstorrent is another step towards this goal. We're glad to have them building alongside us.

Tenstorrent are built for this. An open, full-stack approach to AI compute on RISC-V, designed to be understood and instrumented rather than closed. Tenstorrent's architected openness at the hardware and software layer is what makes runtime assurance possible now, and why they are such an apt partner for us. An open accelerator should not mean inheriting a security gap. The Tenstorrent and Stealthium partnership means the opposite: visibility and threat detection at the accelerator runtime.

Runtime assurance is fast becoming a condition of deploying AI at all, especially in regulated and multi-tenant environments where financial services, telecommunications, and energy operators cannot take isolation on faith. In a multi-accelerator world that assurance must be portable across silicon. It must be a property of the runtime itself.

Vulnerability and risk severity scoring was calibrated for single-tenant infrastructure and orderly patch cycles. Neocloud and multi-accelerator environments break that calibration, because this landscape is frankly young in security maturity. Risks are magnified here. A medium-severity vulnerability in an environment where mutually untrusted tenants share physical hardware will hit like it's high. A low-severity issue on an accelerator with no runtime visibility isn't a low at all. It's an unknown, and unknowns do not carry scores.

Januscape is the clearest example. This was a serious vulnerability for those using nested virtualisation, though thankfully there was a patch organisations could apply and verify. For customers of rent-a-GPU providers and neoclouds that offer access to GPU and AI chips through nested virtualisation, which makes sense given the access you want, Januscape means an attacker can spin up a VM in their own tenant on the shared host, break its isolation and seize control of the machine underneath it, and reach your tenant from there. Training runs, inference workloads, sensitive data and proprietary models, all within reach of an uninvited guest. What was a necessary patch in one environment is a critical risk in an environment you don't own. (Keep in mind: you cannot confirm a host you don't control was ever patched. A patch you cannot verify is not a control.) AI accelerated compute runtime prevention is what closes that distance. For Januscape, Stealthium remains the only publicly documented runtime prevention.

The new multi-cloud and multi-accelerator future is here. The window between a vulnerability being disclosed and an AI infrastructure fleet being patched is measured in weeks. The frameworks that would catch what happens in between mostly do not exist yet.

The question for anyone running AI accelerated compute is not whether their stack can detect the latest attack, published or not. It is what else is executing on their most critical layer that they cannot see.

Hope is not a strategy.


Stealthium's research on GPUBreach, Januscape, and the wider accelerator attack surface is published at stealthium.io/blog.