ABOUT STEALTHIUM
AI is reshaping how the world builds, works, and creates. The systems underneath are becoming some of the most critical infrastructure on the planet. At Stealthium, we work at the edge of accelerated computing, observability, and runtime security. From GPUs to workloads to the fabric connecting them, we're making complex AI environments visible, understandable, and defensible.
WHAT YOU WILL BE WORKING ON
You'll own the reliability of the platform our customers depend on to watch their GPU fleets. That means the ingestion pipelines that never get to drop telemetry, the Kubernetes clusters they run on, and the deployment, alerting, and incident tooling around them. When we say we make AI infrastructure observable, our own infrastructure has to be the proof.
WHAT YOU WILL DO
- Own availability, latency, and capacity for our telemetry ingestion and detection pipelines
- Build and operate our Kubernetes-based infrastructure with infrastructure-as-code
- Design the observability for our own stack: metrics, tracing, alerting, and SLOs
- Automate deployments, rollbacks, and incident response so humans handle only what machines can't
- Run capacity planning for workloads that scale with our customers' GPU fleets
- Take part in a sane on-call rotation and drive postmortems that actually change things
WHAT WE ARE LOOKING FOR
- Strong experience running production systems at scale: Linux, containers, Kubernetes
- Solid grasp of observability tooling — Prometheus/Grafana, OpenTelemetry, or similar
- Experience with infrastructure-as-code (Terraform or similar) and CI/CD pipelines
- Comfort debugging across the stack, from kernel to network to application
- A bias for ownership — you care about the outcome, not just the ticket
NICE TO HAVE (not required)
- Experience operating GPU clusters or other accelerator infrastructure
- Background in security, telemetry-heavy, or data-pipeline products
- Experience with Go and PostgreSQL
WHY JOIN STEALTHIUM
- Work on a problem that matters: securing the infrastructure AI depends on
- Small, sharp team — your work is visible and it ships
- Shape how things are done while the rules are still being written