Expanse unlocks wasted GPU capacity. A daemon runs on every node in your SLURM or Kubernetes cluster and recovers idle compute through three capabilities: resource prediction (right-sizing job submissions before they reach the scheduler), optimisation suggestions (code and config changes researchers can apply themselves), and failure prediction (catching jobs that will fail before they consume hours of GPU time). The waste is enormous and universal across industries. GPU clusters run at 30-40% effective utilisation. Researchers over-request resources by 2-3x as insurance against job failures. Quant funds, AI labs, and national supercomputing centres all live with the same problem. The hard part isn’t monitoring; it’s prediction. Every cluster runs different code, different models, different data. A useful predictor has to learn the cluster, as well as the workload. We do that by ingesting compute history, GPU telemetry, and the source code of submitted jobs to train cluster-specific models that improve over time. We’re four engineers. We ran HPC and GPU training workloads at the largest quant funds and national supercomputing centres. We faced this problem firsthand, and the only fix was to over-provision and burn millions. Ismaeel built the first multimodal HPC resource predictor as research at EPCC (Edinburgh’s Parallel Computing Centre), which beat every published baseline. This is the tool we wish we had.
Expanse
expanse.shLocations
San Francisco, CA, USA
industry
AI Infrastructure · Software
Stage
Pre Seed
founded in
2025
Socials
About
Open jobs at Expanse
This company does not have jobs relevant to this job board at this time.
To view all their jobs, visit their website.