The dashboard looked perfect. H100 GPUs humming along at full tilt, nvidia-smi reporting a pristine 100% utilization. The cloud bill was certainly matching those numbers—every second of premium compute time accounted for.
Then came the uncomfortable truth: those expensive chips were doing next to nothing.
That stark gap sits at the heart of Utilyze, an open-source monitoring tool that MIT-affiliated startup Systalyze released on April 20 with what amounted to a declaration of war against conventional wisdom. "Your GPU dashboard is lying to you," the company announced, arguing that the metrics teams rely on to gauge accelerator performance are fundamentally broken—measuring activity, not actual work.
Consider a benchmark the company published alongside its announcement. Matrix multiplications across three different workload configurations. Standard tooling reported 100% utilization for all three. Utilyze's readings? 2.6%, 32%, and 88% compute throughput, respectively. Hand calculations confirmed Utilyze's numbers within about 2%. The conventional metric stayed pegged at maximum regardless.
It's the kind of discrepancy that makes infrastructure teams queasy.
The Duty Cycle Trap
The problem isn't exactly a secret, though it's rarely front-of-mind when someone checks a monitoring dashboard at 2 a.m. Standard GPU utilization—the number that nvidia-smi and most cloud provider interfaces display—measures something specific and somewhat limited. According to NVIDIA's NVML documentation, it tracks "percent of time over the past sample period during which one or more kernels was executing on the GPU."
Essentially: is anything running? Not: is the silicon actually being pushed?
If a kernel fires up, even one that barely scratches the surface of the chip's arithmetic capabilities, the GPU registers as busy. Memory-bound workloads waiting on data transfers show full utilization while compute units idle. Undersized batch configurations that can't fill the GPU's massive parallelism read as maxed out. The metric answers a yes-or-no question when what teams actually need is something closer to a percentage-of-potential calculation.
Tools like nvtop and rocm-smi inherit this limitation wholesale. They're all plumbing into NVML or equivalent vendor libraries that expose duty cycles, not throughput. Useful for knowing whether a GPU is dead or alive, perhaps less so for understanding whether you're getting what you paid for.
For organizations burning through GPU budgets that rival small country GDPs, the distinction starts to matter.
Going Deeper
Utilyze primarily uses NVIDIA's Perf SDK and CUPTI interface—the same profiling infrastructure that powers tools like Nsight Systems. This requires permissions that standard monitoring doesn't (sudo, CAP_SYS_ADMIN, or privileged containers) and CUDA Toolkit 11.0 or later. In exchange, it surfaces metrics that duty-cycle tools simply cannot report.
Two numbers anchor the tool's dashboard: Compute SOL (speed-of-light) percentage, measuring arithmetic throughput against the GPU's theoretical peak, and Memory SOL percentage, tracking bandwidth across HBM, L2, and L1 caches. For supported configurations, Utilyze also calculates an "attainable" ceiling—a realistic maximum for a given model and hardware setup, acknowledging that few real-world workloads ever hit theoretical peaks.
The catch, at least for now, is scope. Auto-discovery currently works for vLLM inference servers running a specific subset of models on H100-80GB and A100-80GB GPUs within single-node configurations (up to eight GPUs). Support for other inference backends like SGLang appears on the roadmap as "coming soon." The Go binary, linked against CUDA 13.1 by default, runs a WebSocket server on port 8079. MacOS and Windows builds function as remote clients connecting to a Linux profiling host.
One detail worth noting: calculating that attainable ceiling involves sending GPU configuration data to Systalyze's servers. Users uncomfortable with that can disable it via environment variable (UTLZ_DISABLE_METRICS=1), though the documentation could be clearer about what exactly gets transmitted.
Early Days, Early Reactions

The project landed on GitHub on April 22 with a v0.1.0 tag. Three more versions followed within days—v0.1.1 adding high-contrast mode on April 23, then v0.1.2 and v0.1.3 both on April 27. On April 29, the repository had collected 164 stars and 11 forks. Not viral, but respectable for a specialized infrastructure tool.
Systalyze CEO Manya Ghobadi submitted a "Show HN" post. The thread drew 113 points and 28 comments within roughly 24 hours, mostly technical questions about roadmap and feature requests. Maintainer responses confirmed AMD GPU support is planned but unavailable, and that per-process breakdowns—along with traditional metrics like temperature and power—might eventually join the throughput numbers.
An independent review from Pidune AI Insider, published April 28, provided some real-world color. During an H100 fine-tuning task, nvtop reported 98% utilization. Utilyze clocked actual compute throughput at roughly 14%. The reviewer praised the compute-versus-memory breakdown and the attainable ceiling concept while noting the UI remains a work in progress. AMD support, likewise.
GitHub issues hint at typical early-release friction. WSL2 doesn't work due to CUPTI periodic sampler limitations. Windows 11 users have hit "unsupported platform" errors, though Windows is documented to function as a remote client. AMD support has a dedicated tracking issue as of April 28.
What Gets Installed
Setup involves a curl-to-bash script on Linux or macOS, or a PowerShell equivalent for Windows client configurations. First run may prompt an install of CUPTI 12 or later from PyPI. Depending on driver settings, users may need to modify NVreg_RestrictProfilingToAdminUsers=0 to enable non-root profiling access—a step documented in the README but easy to overlook.
The Apache-2.0 license keeps things open. The codebase includes Docker targets and experimental ARM64 builds, though Jetson and Orin support remains theoretical pending broader ARM64 CUDA library availability.
This is emphatically not a drop-in nvtop replacement. No per-process breakdowns yet. No AMD. No consumer GPUs. What it offers is a different angle on performance—one that some infrastructure teams will find clarifying, maybe uncomfortably so.
The Bigger Picture

Systalyze's broader play centers on AI infrastructure optimization. The company positions Utilyze as part of a toolchain for understanding where GPU spending translates to actual computational work—and where it doesn't. LinkedIn data shows the team at somewhere between 2 and 10 employees as of earlier in April, though this hasn't been independently verified.
The measurement problem itself isn't new. NVIDIA's own documentation and conference presentations have acknowledged for years that duty-cycle metrics differ from throughput. Third-party commentary has echoed the same caveat: nvidia-smi reports time busy, not productivity. What Utilyze brings to the table is lightweight, real-time visibility into that distinction without requiring offline profiling sessions or specialized expertise.
Whether that visibility drives infrastructure changes at scale remains an open question. At minimum, it's a sanity check for teams whose dashboards show saturation while costs feel misaligned with output. At maximum, it's a wedge into rethinking how GPU efficiency gets measured and priced—a question with real stakes as accelerator budgets balloon and utilization assumptions underpin capacity planning and vendor negotiations.
The code is public. The benchmarks are reproducible. That gap between 100% and 2.6%—it's sitting there in the announcement post, waiting to be verified or refuted on your own hardware.
Assuming, of course, you're ready for what you might find.
