The new Prime Inference service combines serverless and reserved compute options, leveraging NVIDIA Blackwell GPUs alongside open‑source runtimes such as vLLM and FlashInfer to host frontier open models.
Prime Intellect has unveiled Prime Inference, a new serverless platform that lets developers deploy cutting‑edge open AI models with both on‑demand and reserved compute options.
Key Features of Prime Inference
The service supports NVIDIA Blackwell GPUs, offering the latest hardware acceleration for large language models. It also integrates open‑source runtimes such as vLLM and FlashInfer, enabling low‑latency inference across a variety of model architectures.
Users can choose between a fully serverless mode, which automatically scales resources based on request volume, or a reserved compute mode that guarantees consistent performance for high‑throughput workloads.
Serverless vs. Reserved Compute
In serverless mode, Prime Inference provisions GPU instances on the fly, billing per millisecond of usage. This model is ideal for experimental projects, prototypes, or workloads with unpredictable traffic patterns.
Reserved compute provides dedicated GPU capacity, allowing enterprises to lock in pricing and achieve predictable latency for production‑grade services. Both modes share the same underlying runtime stack, ensuring model compatibility across deployment choices.
Supported Open Models
- LLaMA‑3‑70B
- Mistral‑Nemo‑12B
- Gemma‑2‑27B
- OpenChat‑3.5‑70B
Prime Inference is designed to be model‑agnostic, so developers can bring any compatible open model to the platform without extensive configuration.
Pricing and Availability
The platform launches with a pay‑as‑you‑go pricing tier for serverless usage and a discounted reserved tier for long‑term commitments. Early adopters receive a credit bundle to experiment with the Blackwell GPU fleet.
Prime Inference is now generally available in all regions where Prime Intellect operates, with plans to expand to additional data centers later this year.
MarkTechPost coverage of Prime Intellect launches Prime Inference