📖 The AI Tool Bible
inference

Latency

Time from request to response. For LLMs, split into time-to-first-token (TTFT) and time-per-output-token (TPOT).

Related terms