If your large language model application is already in production and you cannot answer basic questions about latency, token consumption, error rates, or cost per request, you need LLM observability before you ship another feature. The fastest way to get there without adopting a proprietary platform is to instrument your application with a Prometheus client library, scrape those metrics with Prometheus, and visualize them in Grafana. In this guide I walk through exactly that: the metrics worth collecting, the instrumentation code, the Prometheus and Grafana configuration, and the alerting rules that catch problems before your users do.
The appeal of this stack is that it is open source, battle-tested, and already running in most engineering organizations. You do not need a separate LLMOps vendor to get meaningful telemetry. You need to decide what to measure and wire it up correctly.
Why LLM Observability Differs From Standard Service Monitoring
Monitoring a traditional REST service is mostly about request rate, error rate, and duration. LLM workloads add dimensions that standard dashboards ignore, and skipping them leaves you blind to the failure modes that actually hurt.
- Token economics. Cost scales with input and output tokens, not request count. A single verbose prompt can cost more than a thousand cheap ones.
- Variable latency. Time-to-first-token and total generation time behave differently from a typical API call because output length drives duration.
- Upstream provider behavior. If you call a hosted model, rate limits, timeouts, and model-side errors are outside your control but still your problem.
- Quality and safety signals. Refusals, truncated completions, and guardrail triggers are domain-specific events worth counting.
Good LLM observability treats these as first-class metrics rather than log lines you grep after an incident.
The core metrics to collect
Start with a small, high-signal set. You can expand later.
- Request count labeled by model, endpoint, and status.
- Request latency as a histogram so you can compute percentiles.
- Prompt and completion tokens per request.
- Estimated cost derived from token counts and per-model pricing.
- Error count split by cause: rate limit, timeout, provider error, validation error.
- Time to first token if you stream responses.
Instrumenting Your Application for Prometheus LLM Monitoring
The example below uses Python and the official prometheus_client library, but the same metric types exist in the Go, Java, and Node.js clients. The pattern is what matters.
First, define your metrics once at module load:
from prometheus_client import Counter, Histogram
LLM_REQUESTS = Counter(
"llm_requests_total",
"Total LLM requests",
["model", "endpoint", "status"],
)
LLM_LATENCY = Histogram(
"llm_request_duration_seconds",
"LLM request latency in seconds",
["model", "endpoint"],
buckets=(0.25, 0.5, 1, 2, 5, 10, 20, 30, 60),
)
LLM_TOKENS = Counter(
"llm_tokens_total",
"Total tokens processed",
["model", "direction"], # direction: prompt | completion
)
LLM_COST = Counter(
"llm_cost_usd_total",
"Estimated spend in USD",
["model"],
)
Notice the custom histogram buckets. The default Prometheus buckets top out around ten seconds, which is useless for generation that can run a minute. Set buckets that match your real latency distribution, otherwise your p95 and p99 numbers will be meaningless.
Next, wrap the model call. Keep pricing in a lookup table so you can update it without touching call sites:
import time
PRICE_PER_1K = {
"gpt-4o-mini": {"prompt": 0.00015, "completion": 0.0006},
# add your models here
}
def estimate_cost(model, prompt_tokens, completion_tokens):
rates = PRICE_PER_1K.get(model, {"prompt": 0, "completion": 0})
return (
prompt_tokens / 1000 * rates["prompt"]
+ completion_tokens / 1000 * rates["completion"]
)
def call_model(client, model, messages, endpoint="chat"):
start = time.perf_counter()
status = "success"
try:
resp = client.chat.completions.create(model=model, messages=messages)
usage = resp.usage
LLM_TOKENS.labels(model, "prompt").inc(usage.prompt_tokens)
LLM_TOKENS.labels(model, "completion").inc(usage.completion_tokens)
LLM_COST.labels(model).inc(
estimate_cost(model, usage.prompt_tokens, usage.completion_tokens)
)
return resp
except TimeoutError:
status = "timeout"
raise
except Exception:
status = "error"
raise
finally:
LLM_LATENCY.labels(model, endpoint).observe(time.perf_counter() - start)
LLM_REQUESTS.labels(model, endpoint, status).inc()
The finally block guarantees that latency and request count are recorded even when the call fails. That is important: incidents are exactly when you need the numbers, and a naive implementation that only records on success will hide your worst moments.
Exposing the metrics endpoint
Prometheus pulls metrics over HTTP. Expose a /metrics route. If you run FastAPI:
from prometheus_client import make_asgi_app
app.mount("/metrics", make_asgi_app())
For a standalone worker that is not an HTTP server, start the client's own exporter:
from prometheus_client import start_http_server
start_http_server(8001)
A word on cardinality. Do not label metrics with user IDs, full prompts, or request IDs. High-cardinality labels will blow up Prometheus memory and storage. Keep labels to a bounded set: model name, endpoint, status, direction. This is the single most common mistake teams make, and it is painful to unwind later.
Scraping With Prometheus
Point Prometheus at your service. A minimal scrape config:
scrape_configs:
- job_name: "llm-service"
scrape_interval: 15s
static_configs:
- targets: ["llm-service:8000"]
In Kubernetes, prefer kubernetes_sd_configs or the Prometheus Operator's ServiceMonitor so targets are discovered automatically as pods scale.
Confirm data is flowing by querying in the Prometheus expression browser:
sum(rate(llm_requests_total[5m])) by (model, status)
If that returns series, you have working Prometheus LLM monitoring. Everything else is building on top of it.
Building the Grafana LLM Dashboard
Add Prometheus as a Grafana data source, then build panels around the questions you will actually ask during an incident. Here are the queries I put on a first-pass Grafana LLM dashboard.
Request rate by status, to spot error spikes:
sum(rate(llm_requests_total[5m])) by (status)
p95 and p99 latency, the metric users feel:
histogram_quantile(
0.95,
sum(rate(llm_request_duration_seconds_bucket[5m])) by (le, model)
)
Token throughput, split by direction:
sum(rate(llm_tokens_total[5m])) by (direction)
Hourly spend, the number finance will ask about:
sum(increase(llm_cost_usd_total[1h])) by (model)
Error ratio, a clean single-stat for health:
sum(rate(llm_requests_total{status!="success"}[5m]))
/
sum(rate(llm_requests_total[5m]))
Group these into rows: one for traffic and errors, one for latency, one for cost and tokens. Keep the top of the dashboard answerable at a glance. Nobody reads the twentieth panel during a page.
Alerting on What Actually Breaks
Dashboards are for investigation. Alerts are for waking the right person. Define Prometheus alerting rules for the conditions that correlate with real user pain and real cost overruns.
groups:
- name: llm-observability
rules:
- alert: LLMHighErrorRate
expr: |
sum(rate(llm_requests_total{status!="success"}[5m]))
/ sum(rate(llm_requests_total[5m])) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: "LLM error rate above 5% for 10 minutes"
- alert: LLMLatencyRegression
expr: |
histogram_quantile(0.95,
sum(rate(llm_request_duration_seconds_bucket[5m])) by (le)
) > 15
for: 10m
labels:
severity: warning
- alert: LLMSpendSpike
expr: sum(increase(llm_cost_usd_total[1h])) > 50
for: 5m
labels:
severity: warning
Tune the thresholds to your baseline rather than copying mine. The for clause matters as much as the expression: it prevents a single scrape blip from paging someone at 3 a.m. Start with conservative thresholds and tighten them as you learn your normal range.
Extending Beyond the Basics
Once the foundation is solid, there are natural next steps for a mature LLM telemetry practice.
- Time to first token for streaming endpoints, recorded as its own histogram. Perceived responsiveness lives here.
- Guardrail and refusal counters so you can see when safety filters or validation logic reject output.
- Cache hit rate if you use a prompt or response cache, since it directly offsets cost.
- Retrieval metrics for RAG systems: documents retrieved, retrieval latency, and reranker scores.
- Correlation with traces. Prometheus answers "what" and "how much." Pair it with distributed tracing when you need "where" and "why" for a specific slow request.
When you are deciding how far to invest, weigh your build-versus-buy tradeoffs the way you would for any platform component. The Prometheus and Grafana combination covers metrics and alerting extremely well. It is not a tracing system or an evaluation harness, and that is fine. Use the right tool per concern rather than forcing one stack to do everything.
If you want help standing up this kind of instrumentation across services, our data and MLOps capabilities cover platform engineering for AI workloads end to end. Teams in regulated or high-throughput settings often have additional requirements around retention and auditability, which we address in our industry-specific engagements.
FAQ
What is LLM observability?
LLM observability is the practice of collecting and analyzing telemetry from large language model applications: latency, token usage, cost, error rates, and quality or safety signals. It extends standard service monitoring with dimensions unique to generative workloads, so you can diagnose problems and control spend in production.
Can I use Prometheus and Grafana for LLM monitoring without a dedicated LLMOps tool?
Yes. Prometheus and Grafana handle metrics collection, storage, visualization, and alerting, which covers the core of LLM observability. Dedicated LLMOps tools add capabilities like prompt-level tracing and automated evaluation, but you can get strong production coverage with the open-source stack first and add specialized tooling only where you have a concrete gap.
How do I avoid high cardinality in LLM metrics?
Label metrics only with bounded, low-cardinality dimensions such as model name, endpoint, status, and token direction. Never attach user IDs, request IDs, or full prompt text as labels. If you need per-request detail, put it in logs or traces, not in Prometheus metric labels.
How should I track LLM cost in Prometheus?
Maintain a per-model pricing table in your application, compute estimated cost from the prompt and completion token counts returned by the provider, and increment a llm_cost_usd_total counter. You can then query hourly or daily spend in Grafana and alert on spending spikes before the invoice surprises you.



