Latency
Definition
The time elapsed between sending an API request and receiving the first generated token (Time to First Token).
Example Case
Groq delivering outputs with a latency of 150ms.
The time elapsed between sending an API request and receiving the first generated token (Time to First Token).