News confirmed medium confidence

Together AI Adds IBM Cloud B300 Cluster for Open-Model Inference

The dedicated infrastructure assigns the inference layer to Together AI, the cloud to IBM and B300 GPUs plus Spectrum-X networking to NVIDIA, but leaves capacity and performance undisclosed.

Edited by Tyronne Panaino

Together AI announced on October 6 that it is expanding enterprise inference capacity through a dedicated cluster of NVIDIA B300 GPUs on IBM Cloud, backed by NVIDIA Spectrum-X Ethernet networking. Together AI will operate the inference layer, IBM will provide the cloud infrastructure, and NVIDIA will provide the GPUs and networking. For organizations evaluating open-model inference, the announcement identifies a new infrastructure configuration, but it does not establish lower latency, lower cost or greater reliability.

Together AI describes itself as the first customer running on what it calls the first dedicated, large-scale inference cluster of its kind on IBM Cloud. That description comes from Together AI rather than an independent assessment. The announcement does not identify other users, a deployment region, a date for broader customer access, pricing or service-level terms.

A three-layer operating model

The clearest part of the release is its division of responsibilities. Together AI operates the software layer that serves open models. IBM supplies the cloud environment in which the dedicated cluster runs. NVIDIA supplies the B300 accelerators and Spectrum-X Ethernet networking. The source therefore supports a three-layer architecture claim; it does not support treating all three companies as jointly guaranteeing one measured performance result.

Together AI calls the cluster purpose-built for inference. In this announcement, dedicated describes the infrastructure arrangement, but the company does not define its isolation model, reservation terms, data-residency options or service-level agreement. It also gives no GPU count, network topology, supported-model list or regional footprint for the new capacity.

Why Together AI says it needs more capacity

Together AI reports that its platform serves hundreds of trillions of tokens per month to more than one million developers. That first-party scale claim explains why the company is adding a dedicated inference cluster, but it is not an audited utilization figure and does not show how much of that traffic will move to IBM Cloud.

The company positions the collaboration as infrastructure for running open models in production. The material change is the named combination of Together AI's inference platform, IBM Cloud, NVIDIA B300 GPUs and Spectrum-X networking. The release does not introduce a new model, API or benchmark, and it does not say that every Together AI customer is already using the cluster.

What buyers can and cannot conclude

The announcement supports a limited conclusion: Together AI says it is running as the first customer on a dedicated B300 cluster on IBM Cloud and has assigned the three infrastructure layers to named providers. It does not provide measured throughput, time to first token, end-to-end latency, utilization, energy use, failure rates, security-test results or price comparisons.

Without those measurements, readers cannot verify Together AI's broader claims about efficiency, reliability or enterprise readiness from this release alone. The next useful checkpoints would be an IBM or NVIDIA technical confirmation, disclosed cluster specifications, a customer-availability notice, independently reproducible benchmarks and evidence from production users.

Status

Confirmed. Together AI published the infrastructure announcement on its official site. Internal confidence is medium because the central deployment details come from one participant and have not been independently corroborated in the evidence reviewed for this article.

Sources

Update note: Last reviewed 2026-10-10. We will revise this post if IBM, NVIDIA or Together AI publishes specifications, availability details or measured results.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More News coverage