BIP America News & Media Platform

collapse
Home / Daily News Analysis / IBM, Together AI team to offer open-source AI inference in the cloud

IBM, Together AI team to offer open-source AI inference in the cloud

Sep 07, 2026  Twila Rosenbaum  1 views
IBM, Together AI team to offer open-source AI inference in the cloud

IBM and Together AI have announced a $240 million agreement to deliver open-source AI inference from IBM Cloud. Together AI, a specialist neocloud provider, will use a large cluster of Nvidia HGX B300 systems connected through Nvidia Spectrum-X Ethernet networking. IBM will own and operate the infrastructure, while Together AI will run its open-source model serving platform on top for enterprise customers.

Together AI’s platform supports widely used open models, including DeepSeek, Nemotron, MiniMax, Kimi, and GLM. The company says these models give developers the freedom to customize and fine-tune lower-cost alternatives to proprietary AI systems. IBM and Together AI expect that the combination of IBM’s cloud enterprise capabilities, Nvidia’s high-performance GPUs, and Together AI’s inference software will make it easier for businesses to run production-grade open-source AI workloads.

Why the IBM-Together AI partnership matters

The deal is a significant bet on open-source AI inference. It also demonstrates how IaaS, or infrastructure-as-a-service, is evolving to meet AI-specific requirements. Traditional public cloud IaaS was built for general-purpose enterprise applications. AI inference, by contrast, needs specialized accelerators, low-latency networking, large memory pools, and software that can schedule models dynamically.

IBM says the hybrid environment will give organizations a reliable foundation to build, deploy, and scale AI systems. The technology behind the agreement also depends on infrastructure that IBM is co-developing with Nvidia. The two companies said earlier that they would expand their collaboration to grow the use of GPU-native data analytics, intelligent document processing, and on-premises or regulated infrastructure deployments. All of these efforts are aimed at moving AI from pilot to production.

Nvidia HGX B300 and Spectrum-X in the cloud

At the heart of the service will be Nvidia HGX B300 systems. These are high-density server platforms optimized for both training and inference workloads, offering large accelerator memory and high-throughput performance. They are designed to run large language models at scale, including open-weights models that require significant computational capacity.

The cluster will also use Nvidia Spectrum-X Ethernet networking. Spectrum-X is designed to enable predictable performance and low latency for AI workloads on Ethernet networks. That is important for inference, especially when a single user request may require routing to different nodes or when multiple reasoning steps happen in rapid succession. Nvidia Spectrum-X can reduce packet conflicts, provide fair bandwidth between jobs, and improve utilization for constantly running inference services.

Together AI’s cloud native inference engine sits above that infrastructure.


Source: Network World News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy