Service health
Request rate, queue time, time to first token, total latency, errors, retries, and saturation.
Monitor service health, model behavior, output quality, and spend together so that a healthy API cannot conceal a degraded product experience.
Request rate, queue time, time to first token, total latency, errors, retries, and saturation.
Model and prompt version, token counts, finish reason, tool calls, fallbacks, and context truncation.
Task success, groundedness, safety checks, user feedback, and reviewed production samples.
Requests and tokens by feature, provider charges, cache effectiveness, quota use, and cost per successful task.