Executive Summary
AI Gateway has launched a new feature that displays live performance metrics, specifically throughput and latency, for hundreds of available AI models. Updated hourly based on real customer traffic, this data is integrated into the model list, detailed model pages for provider comparison, and is also accessible programmatically via a REST API. This enhancement is designed to help users make informed decisions when selecting the best-performing model and provider for their application's needs.
Key Takeaways
* Live Performance Data: The feature introduces two key metrics: latency (time-to-first-token) and throughput (tokens per second).
* Hourly Updates: All metrics are refreshed every hour, reflecting recent performance based on live AI Gateway traffic.
* Three Points of Access:
* Model List: Main list now includes sortable columns showing the best P50 latency and throughput for each model across all its providers.
* Model Detail Pages: Allows for direct comparison of P50 performance metrics between different providers for the same model.
* REST API: Provides programmatic access to P50 and P95 latency and throughput data for integration into automated workflows.
* Informed Decision-Making: The primary goal is to empower developers to choose the optimal model by providing transparent, objective performance data.
Strategic Importance
This update positions AI Gateway as a more intelligent routing and discovery tool, moving beyond simple access to enabling performance-based optimization for developers building on LLMs.