Live Provider Telemetry
API Endpoint Latency & Speed Matrix
Real-time benchmark tracking Time To First Token (TTFT), throughput (Tokens/Sec), and 24-hour SLA uptime across major cloud inference engines.
Auto-refreshing • Last updated: Just now
Fastest Throughput
485.0Tokens / sec
Groq LPU Inference EngineLowest Latency (TTFT)
65 msTime To First Token
Groq / Gemini 2.0 FlashGlobal Network Uptime
99.96%24h SLA
All Endpoints Operational| Provider / Engine | Region | Status | Latency (TTFT) | Throughput (Tokens/sec) | 24h Uptime |
|---|---|---|---|---|---|
| Groq LPU Engine | US-West (Oregon) | Operational | 65 ms TTFT 35ms total latency | 485 t/s | 99.98% |
| Google Gemini 2.0 Flash | Global (Google Cloud) | Operational | 220 ms TTFT 110ms total latency | 115 t/s | 99.99% |
| Together AI Open Models | US-East & EU-Central | Operational | 240 ms TTFT 125ms total latency | 135.2 t/s | 99.89% |
| OpenAI GPT-4o Endpoints | US-East (N. Virginia) | Operational | 280 ms TTFT 142ms total latency | 84.5 t/s | 99.98% |
| Mistral AI Platform | EU-West (Paris) | Operational | 295 ms TTFT 155ms total latency | 72.8 t/s | 99.96% |
| Anthropic Claude API | US-East (AWS us-east-1) | Operational | 310 ms TTFT 165ms total latency | 78.2 t/s | 99.95% |
| DeepSeek API Endpoint | Asia-East & US-West | Operational | 340 ms TTFT 195ms total latency | 92.4 t/s | 99.9% |