IBM partners with Together AI to deploy a large‑scale inference cluster on IBM Cloud using NVIDIA HGX B300 systems, aiming to accelerate open‑source AI workloads for enterprises.

IBM and Together AI have entered a multi‑year partnership to build a large‑scale inference cluster on IBM Cloud, leveraging NVIDIA HGX B300 systems. The collaboration aims to accelerate open‑source AI workloads for enterprises, providing a flexible, high‑performance environment for models such as LLaMA, Falcon and Stable Diffusion.

Partnership Overview

The agreement combines IBM’s hybrid cloud expertise with Together AI’s open‑source model development and NVIDIA’s cutting‑edge AI infrastructure. Together AI will run its inference services on IBM Cloud, while IBM will integrate NVIDIA HGX B300 accelerators to deliver the necessary compute power.

Technical Architecture

The inference cluster will be built on IBM Cloud’s dedicated bare‑metal servers equipped with NVIDIA HGX B300 GPUs, each offering up to 8 × NVIDIA Hopper Tensor Cores. This setup enables high‑throughput, low‑latency serving of large language models and generative AI applications.

  • IBM Cloud provides secure, enterprise‑grade networking and data governance.
  • NVIDIA HGX B300 delivers up to 2 petaflops of AI performance per node.
  • Together AI supplies optimized model libraries and inference APIs.

Benefits for Enterprises

Enterprises can now access scalable, cost‑effective inference for open‑source models without managing on‑prem hardware. The solution promises faster time‑to‑value for AI initiatives, reduced operational complexity, and compliance with data residency requirements.

Future Roadmap

IBM and Together AI plan to expand the offering with additional model support, automated scaling features, and integration with IBM’s Watsonx AI platform. Ongoing collaboration with NVIDIA will ensure the infrastructure stays at the forefront of AI hardware advancements.

IBM newsroom coverage of IBM and Together AI partnership