🛠️ Tools / /via The Verge / updated Jul 14, 2026

OpenAI and Microsoft Release Triton 3 Inference Framework

OpenAI and Microsoft launched Triton 3 on July 12 2026 with automatic kernel fusion and 3x faster MoE inference. The open-source framework supports PyTorch 2.7 and CUDA 13. Developers can achieve 420 tokens per second on H200 GPUs for 70B models.

#OpenAI#Microsoft
~/ Tools/ OpenAI and Microsoft Release Triton 3 Inference...

OpenAI and Microsoft released Triton 3 on July 12 2026. The update introduces automatic kernel fusion for mixture-of-experts models and delivers up to 3x faster inference than the prior version. It supports PyTorch 2.7 and CUDA 13 out of the box.

Benchmark tests show 420 tokens per second on a single H200 GPU for a 70 billion parameter model. Memory usage dropped 35 percent through new paging techniques. The framework is available on GitHub under the same permissive license as previous releases.

Triton originated at OpenAI in 2021 and was open-sourced in 2022. Microsoft contributed major performance improvements through its Azure AI team. Over 180,000 developers currently use the library according to GitHub metrics.

The new version adds native support for FP8 and INT4 quantization plus dynamic batching for production serving. OpenAI used Triton 3 internally to power the recent o3 model release. Microsoft integrated it into Azure Machine Learning endpoints the same day.

Why this matters

Unified tooling reduces fragmentation between research and production environments. Smaller teams gain access to optimizations previously limited to hyperscalers. The collaboration signals deeper technical integration between OpenAI and Microsoft despite ongoing antitrust scrutiny.

Future releases will target AMD and Intel GPUs by early 2027. The maintainers committed to quarterly updates with community governance starting in September 2026.

share
𝕏 FB
← cd ../news